The short answer, if you only want one: Perplexity when you need sources you can click and check, Claude when the evidence disagrees with itself and someone has to weigh it, Gemini when you want a long report quickly, and ChatGPT when the output has to read like a finished document. Most people who use these seriously end up paying for two.
The longer answer is the part that matters, and it has little to do with which logo you pick. A large audit of deep research agents found link validity above 94% and factual accuracy between 39% and 77%. The links work. What the report claims those links contain is a separate question, and it goes wrong often enough that reading the citations is now part of the job.
So this piece has two halves. A current read on which tool suits which kind of work, and then the half that gets published far less often: how to check a research report in about ten minutes before you forward it to anyone who matters.
What deep research actually does
A normal AI answer searches, finds something that looks right, and stops. Deep research works differently in one specific way: it runs a loop. It plans, searches, reads, notices what is missing, searches again, and only then writes. OpenAI's version typically runs ten to thirty minutes and consults somewhere between fifty and two hundred sources.
That loop is why the output feels like a research analyst produced it. Headings, a summary, caveats, a reference list. It is also why the failure mode is unusual. A short wrong answer looks wrong on sight. A twelve-page report with forty citations looks like work, and it borrows the authority of its format whether or not it earned it.
Understanding the loop also tells you where the errors enter. Each stage compounds: a search that misses the best source, a page read partially, a summary that drops a qualifier, a synthesis that smooths two findings into one sentence. By the time it reaches the draft, the report has no memory of which of its claims were solid and which were assembled from fragments. Everything comes out in the same confident register.
The gap between a citation and a fact
This is the finding that should change how you read any of these tools.
In audits of deep research agents, links resolve and sources look relevant far more often than the claims attached to them hold up. Link validity above 94%. Source relevance above 80%. Factual accuracy between 39% and 77%, depending on the model and the subject. A citation can point at a real, relevant, respectable source and still misrepresent what that source says.
That is a harder problem than invented references, because an invented reference announces itself the moment you click. A real link under a wrong claim stays quiet.
Invented references are getting more common too, and the trend in academic publishing is worth seeing because it tells you what happens when verification gets skipped at scale. A study published in The Lancet on 7 May 2026, led by Maxim Topaz at Columbia, examined over two million papers and 97 million citations and found the rate of fabricated references climbing steeply:
- 2023: 1 in 2,828 papers contained at least one fabricated reference
- 2025: 1 in 458
- First seven weeks of 2026: 1 in 277
These are peer-reviewed papers, written by researchers, reviewed by other researchers, published in journals with editorial staff. Survey work suggests roughly three quarters of peer reviewers do not thoroughly check the reference lists they are reviewing. If citations slip through that process, they will slip through a report you skimmed on a Tuesday afternoon.
Published hallucination rates for citation tasks vary enormously, from roughly 11% to over 50% across commercially deployed systems in one 2026 cross-model audit, and wider still in studies that include smaller models. Treat any single headline number with suspicion, including these. The conclusion that survives is directional: the rate is high enough, and varies enough by topic, that checking is not optional on anything consequential.
Which tool for which job
Reviewed October 2026, drawn from published capabilities and third-party testing. I have not re-run a controlled head-to-head since the one in Letter 25, and that was March 2025, so read this as a map and not a verdict. Limits and prices move constantly; check the current plan before you buy.
- Perplexity. Built around cited search from the beginning, and still the quickest route to sources you can open and check. The strongest default when the citations are the point, and the easiest to audit because the trail stays visible as you read.
- Claude. The best of the group at holding contradictory evidence in view instead of smoothing it into a clean answer, and at stating plainly that two sources disagree. Also the one to reach for when the material is a set of documents you upload yourself.
- Gemini. Fast at an acceptable standard, with a large context window and the natural choice if your material already lives in Google Workspace.
- ChatGPT. The longest and most polished reports, and the strongest when the output needs to be presentable without a rewrite. The polish is also what makes it the easiest to over-trust.
Two practical warnings before you commit money.
First, deep research runs are metered, and the allowances have moved sharply. Reported figures for 2026 put ChatGPT in the region of ten runs a month on the entry paid tier, with substantially more on the expensive ones, and Perplexity cut its Deep Research allowance during early 2026. If you plan to lean on this daily, confirm the current number on the plan page, because it is the detail every comparison article gets out of date first.
Second, the quality difference between these tools is smaller than the quality difference between a vague prompt and a specific one. Which brings us to the part you control.
How to ask for a report you can actually use
Most disappointing deep research output traces back to a question that was too broad to answer well. Four things tighten it.
- Name the decision. "Should we move our support triage to an open-weight model this quarter" produces a usable report. "Tell me about open-weight models" produces an encyclopedia entry.
- Set the time boundary explicitly. Say "published in the last twelve months" and say what today's date is. Models carry stale material forward with complete confidence unless you fence the window.
- Ask for the disagreement. Add "where do credible sources disagree, and what is the strongest argument against the conclusion". This single line changes the shape of the output more than any other instruction.
- Demand source types. "Prefer primary sources, regulatory text and peer-reviewed work; mark anything that comes from a vendor." You are giving yourself a verification map inside the report.
One more that costs nothing: ask it to flag its own weakest claims. Models are surprisingly willing to tell you which parts of a report they are least sure about, and that list is a good place to start checking.
The ten-minute check
Here is the pass I run before a research report goes anywhere that matters.
- Pick the three claims the decision rests on. Not the whole report. The sentences that would change what you do. Everything else can wait.
- Open the source for each one. The source itself, not the report's summary of it. This single step catches the majority of problems, because it separates "the link works" from "the link says this".
- Check the number and the date together. Figures get carried forward correctly from the wrong year constantly. If a number matters, find it in the original and confirm which period it describes.
- Look for the quiet qualifier. Source says "in a sample of 40 undergraduates", report says "studies show". That compression is where most false confidence enters.
- Ask what is missing. A report returns what it found. Search once yourself for the strongest argument against the conclusion. If the report never mentions it, you have learned something about the report.
- Check who benefits. Vendor blog, trade body, or an independent source. All three can be right, and they fail in different directions.
If you only ever do the second step, you will still be ahead of most people using these tools.
What that looks like in practice
Say a report tells you that a particular approach "cuts support resolution time by 40%, according to industry research", with a citation.
You open the link. It is a real consultancy study, published eighteen months ago. Inside, the 40% figure describes a pilot at three companies, all of them in the same sector, measuring first-response time rather than resolution time, and the study's own authors call it preliminary. None of that is dishonest on the report's part. Every piece of it is a small compression, and together they turn a narrow pilot result into an industry fact.
That is the typical case. Not a fabrication, a flattening. You catch it by opening one link, and it takes ninety seconds.
What a citation actually proves
Worth being precise, because the word carries more weight than it has earned.
A citation proves that a document exists at a URL. It is evidence about the report's process and very little evidence about the claim. It does not establish that the document says what the sentence says, that the document is any good, that the figure was read correctly, or that a contradicting source was considered and rejected instead of never found.
The reason this matters more with deep research than with a normal answer is volume. Forty citations are impossible to spot-check casually, so most readers check zero. The format defeats the instinct it was built to satisfy.
Three failure modes worth knowing by name
- The confident orphan. A claim with a real, relevant citation that does not contain it, usually because the model inferred the claim from material surrounding the quote. The hardest kind to catch and the most common.
- The stale number. A correct figure from a superseded source, presented in the present tense. Pricing, market size, model capabilities and regulation all rot fast, and 2024 numbers are still circulating as current.
- The agreeable sweep. A literature that genuinely disagrees, reported as settled. Usually a sign the tool found three sources from the same side of an argument and stopped looking.
When to skip deep research entirely
It is the wrong instrument more often than the marketing suggests. Four cases where a normal search or a phone call beats it.
One authoritative source already answers the question. A report built around a fact you could look up adds length, delay and risk without adding knowledge.
The answer has to be exactly right. If you would need to verify every line anyway, the checking costs more than the research saved. Go to the primary source first.
The subject moves faster than the index. Model pricing, API limits and anything regulatory change between the crawl and your reading of it. Go to the vendor's own page.
You already know the answer and want support for it. This is the use that feels best and serves you worst, because a tool that searches until it finds agreement will always find some.
Common questions
Which AI is best for deep research in 2026?
There is no single winner, and the honest split is by job. Perplexity for sources you can check quickly, Claude for weighing evidence that disagrees, Gemini for speed and for material already in Google Workspace, ChatGPT for the most polished long report. Many heavy users pay for two, because the switching cost is low and the strengths genuinely differ.
Can you trust AI deep research reports?
Not without checking the claims your decision rests on. Audits of deep research agents found link validity above 94% but factual accuracy between 39% and 77%, which means citations frequently point at real sources that do not support the sentence attached to them. Treat the report as a well-organised starting point, and verify the two or three claims that actually drive the outcome.
Do AI tools make up citations?
Yes, and the rate varies widely by model and topic, from roughly 11% to over 50% across deployed systems in one 2026 audit. Fabricated references are also rising in published academic papers, from 1 in 2,828 in 2023 to 1 in 277 in early 2026 according to a Lancet study covering two million papers. Invented references are the easier problem to catch; real links under wrong claims are harder and more common.
How do you fact-check an AI research report?
Pick the three claims your decision depends on, open the original source for each instead of the report's summary, check the number and its date together, watch for qualifiers that got compressed away, search once for the strongest counter-argument, and note who benefits from each source. Ten minutes covers most of the risk.
Is deep research worth paying for?
It is worth it if you regularly need a landscape across many sources and you will actually verify what comes back. Runs are metered and the allowances moved during 2026, so check the current limit on the plan you are considering. If your questions usually have one authoritative answer, a standard subscription is enough.
What is the difference between deep research and normal AI search?
Normal search finds a few good-enough sources and stops. Deep research runs a loop: it plans, reads, notices gaps, searches again and then writes, typically over ten to thirty minutes and across fifty to two hundred sources. That produces something shaped like an analyst's report, which is useful and also more persuasive than its accuracy warrants.
Which deep research tool has the best citations?
Perplexity is generally the easiest to audit, because the citation trail stays visible while you read and the sources sit next to the claims. Easiest to audit is a different property from most accurate, though. Whichever tool you use, the check is the same: open the source for the claims that matter and confirm it says what the report says it says.
What I would do
Use the tool that fits the job and stop hunting for the one that wins everything, because that comparison has no stable answer and the time is better spent elsewhere. Then build the ten-minute check into how your team works, as a habit rather than a policy document, and make the second step non-negotiable: open the source, not the summary.
The teams that get real value from deep research treat its output as a well-organised draft produced by a capable assistant who has never once been held accountable for being wrong. That is the whole discipline, and it generalises far beyond research reports. If you want the full method, including how to judge work you did not produce and how to decide what evidence a claim actually needs before you act on it, that is what the Verification Masterclass covers.
Charafeddine Mouzouni




