Marvin Labs
Source Judgment: What Makes an AI Research Tool Reliable
Live with Marvin

Source Judgment: What Makes an AI Research Tool Reliable

8 min readJames Yerkess, Senior Strategic Advisor

The models behind AI research tools are now good at answering from the material in front of them. In a May 2026 study, the best of six commercial chatbots answered more than 90% of multiple-choice questions correctly about news events reported only hours earlier. Where they failed, Suzgun et al. traced more than 70% of the errors to retrieval: the model worked from the wrong material or missed the right one.

That puts the weight on source judgment. Deciding which sources to use, how much each one deserves, and when to distrust one is what analysts do before they write a number into a model. An AI research tool has to do the same thing on every question, across every source it can reach, and that is now what separates one tool from another.

This piece builds on a Marvin Labs LinkedIn Live discussion with Alex Hoffmann (Co-Founder & CEO, Marvin Labs) and moderator James Yerkess (Former Global Head of Transaction Banking & FX, HSBC Wealth Management) on Marvin 2.0, the AI behind AI Analyst Chat and Deep Research Agents, and the design choices behind it.

Source judgment starts with a hierarchy

Marvin 2.0 starts every answer from primary financial content: the filings, earnings calls, and press releases that Automated Data Import ingests and validates. When a question can be answered from those, it is answered from those.

Beyond them sits the connector catalog, more than 30 sources covering fundamentals, estimates, expert networks, alternative data, and a firm's own notes. Public data such as FRED macro series and EIA energy statistics is hosted by Marvin Labs and on from the first day. The rest connect with the analyst's own provider account through a sign-in or one key, and a source the catalog does not list, including a system the firm runs itself, can be added by hand.

Chat and agents reach for a connected source when the question needs data the documents do not hold, and a connected source does not override what the company itself disclosed. The hierarchy is the first layer of judgment: first-party disclosure, then the data an analyst has chosen to connect, then everything else.

Judgment has to work on sources nobody has seen

A whitelist does not scale to this. Marvin Labs cannot know in advance which providers a client will connect or which documents they will upload, so the agent has to assess each source as it arrives. It validates what it retrieves and treats third-party data with suspicion until it checks out.

Take a question about the supply and demand balance for oil leaving the Middle East through the Strait of Hormuz. Marvin Labs holds no oil data of its own, so the answer comes from agency statistics the agent finds and evaluates, and it has to pass over the blog posts, social media threads, and opinion pieces that rank alongside them. The same applies to any widely followed company. A question about sentiment on Manchester United's finances will surface far more fan opinion than first-party disclosure, and a useful answer weighs the club's own accounts above a rival supporter's view.

General-purpose chatbots show what happens without that discipline. In NewsGuard's one-year audit, published in September 2025, the ten leading chatbots stopped declining questions altogether, and the share of answers repeating false claims rose from 18% to 35%.

The new source of hallucinations right now is really you have bad data, and the bad data gets parroted as if it's good data.

— Alex Hoffmann

Your own documents get the same judgment

Analysts bring their own material in two ways. A chat attachment, whether a PDF, text, Markdown, CSV, or image, is read in that conversation. Keep it when the chat asks, and it joins the company's document list with a title, a date, and a type such as broker research, internal research, or a third-party presentation, where later chats and agents find it without being told. Private document upload does the same without opening a chat.

The classification is source judgment at work. An internal investment thesis, a sell-side note, and a consultant's market study each carry different authority, and the agent reads them that way. A rough five-year guess at Tesla's revenue in a spreadsheet gets less weight than a fully built, well-argued model, even though both come from the analyst's own drive.

It's able to get a sense for the quality of your inputs as well.

— Alex Hoffmann

The structure around the document matters as much as the weighting. Yerkess ran his own analysis of Manchester United through it, weighing the Glazers' apparent appetite to sell against INEOS's longer-term investment pattern, and got back an analysis organized around the questions an analyst would ask of the material. In his account, the same file in a general-purpose chatbot produced a weaker result.

Evaluation builds the judgment, citations show it

Judgment of this kind comes from evaluation. Marvin 2.0 came out of a loop of real research questions, the failure cases they surface, and repeated changes to the agent until those cases pass. The release notes list what that work produced, including financial lookups that tell an adjusted margin from an as-reported one and a single ranked search across filings, transcripts, and uploaded documents. The loop continues after each release.

Citations make the result visible. Each answer cites the documents it read as numbered links and ends with a list of its sources, so an analyst can see which sources the agent trusted and open any of them before a figure goes into a model or a deck. Tables and charts carry the same attribution, so a figure keeps its source when it moves into a deck or a memo.

Comparisons keep each company's sources separate

A comparison across companies tests source judgment hardest, because it mixes the fiscal calendars, reporting currencies, and adjusted measures of several issuers in a single answer.

Marvin 2.0 keeps the companies apart. Ask for the revenue of Google and Apple in one question and each company gets its own agent, which works from that company's sources and cannot see the other's. An orchestrating agent above them works out what each one needs to answer, collects the results, and does the comparison. Follow-up questions go back to the relevant company agent, so every figure in the final answer traces to the right issuer's documents.

AI Analyst Chat comparing the most recent annual gross margin for Apple, Microsoft, and SAP in a source-attributed table
A three-company comparison in AI Analyst Chat. Each company's analyst works under its own logo, SAP's as-reported and adjusted figures sit side by side, and every company stays in its reporting currency.

Macro questions for equity analysts

The same hosted sources made standalone macro questions possible. Macro data first went in so analysts could ask company questions with a macro edge, such as how much revenue grew above inflation or GDP. Reviews of anonymized chats then showed analysts starting from macro questions with no company attached. With vetted sources already connected, answering them directly took one more pass through the evaluation loop.

The scope is deliberate: macro research for equity analysts, grounded in official statistics. Specialist macro economists will still want their own tools.

Scheduled agents apply the same judgment on a trigger

Deep Research Agents write reports in a format Marvin Labs configures and quality-manages, with the same source hierarchy as Chat. Scheduled agents run them without a click, at a set time or when a company in scope publishes a new filing, an earnings press release, an earnings call transcript, or the earnings filing itself. Triggers on market moves, such as a report when a stock falls 5% overnight, are in development. Scheduled runs need a paid plan. Manual runs are included in the free evaluation.

How to test source judgment in any AI research tool

  • Ask a question where the obvious web answer is wrong. Check whether the answer follows the top search result or the primary source.
  • Ask about a widely discussed company. See whether opinion pieces and forum posts appear among the cited sources.
  • Upload a rough estimate and a full model for the same company. See which one shapes the answer.
  • Open the citations. Every figure should trace to a document you can read, at the page or passage it came from.
  • Compare two companies with different fiscal years or currencies. Check that adjusted and as-reported figures stay labeled and each company stays in its reporting currency.

AI Analyst Chat is a good place to run all five.

For the full conversation, including an exchange on telling poorly thought-out management apart from poorly thought-out analysis, watch the video above.

Run this analysis on your own names

Put the same questions to the filings, earnings calls, and financials of the companies you cover, with every answer cited to the source. Free to evaluate on 15+ companies.

Explore the companies

Discuss your coverage and workflow with the founder in a 30-minute call.

James Yerkess

by James Yerkess

James is a Senior Strategic Advisor to Marvin Labs. He spent 10 years at HSBC, most recently as Global Head of Transaction Banking & FX. He served as an executive member responsible for the launch of two UK neo banks.