Citing Is Not Verifying: One Prompt, Two Systems, and the Layer Beneath
Perplexity Deep Research and a multi-stage verification pipeline were given the identical prompt. Both delivered real sources – and both produced the same class of errors. The difference is not a better model, but a separate verification step that also runs against itself.
Framing note: This article compares a single Perplexity Deep Research run with a single pass of our own research and verification pipeline – same prompt, same topic. It is a case study, not a study. Where external research supports the pattern or contradicts it, it appears in the text.
Superseded, and a correction (30 August 2026): This article has been replaced by Three Systems, One Prompt, which adds a third research tool. It stays online here as a record. While working on the successor, our own evidence check found an error in the literature-search paragraph that it had missed on the first pass: two different measurements from the same study had been conflated. The paragraph below is corrected.
Dieser Artikel ist auch auf Deutsch verfügbar: Zitieren ist nicht Belegen
Evidence on file: the unabridged Perplexity report verbatim, and the article our own pipeline produced, which this text discusses.
One Prompt, Two Systems
On 29 August 2026 we ran an experiment. Two systems received the identical brief, word for word:
“Analyze the implications of eating meat for health, animal wellbeing, environmental impact. Come up with the best strategy to achieve the best compromise along those three axes.”
System one: Perplexity Deep Research, an automated research system that runs dozens of searches on its own, reads sources, and builds a report with citations from them. The result: roughly 1,400 words with 32 sources, finished in a few minutes. One run on a Pro subscription, Deep Research mode, no post-editing; we did not document which model actually served the run – more on that later.1
System two: our own article pipeline. Five independent research agents – separately running AI instances that know nothing of each other – then produce an outline and a draft, which run through a six-stage verification chain. The result: 2,763 words with 35 footnotes, every load-bearing claim checked against the raw text of its source – with four documented exceptions, recorded as such in the verification log.2 The effort: around 20 agent runs, about 1.7 million agent tokens – a token, the unit in which AI text is processed and billed, corresponds roughly to a word fragment – and about one working day.3
The yardstick for this comparison is not ours. Perplexity itself promises, in its launch post, “expert-level analysis across a range of complex subject matters” and completion of most research tasks “in under 3 minutes”.1 So we measure the product against its own claim – and, as will become clear, our pipeline along with it.
Because the thesis of this article is: Citing is not verifying. Deep-research tools deliver real sources and fail one layer deeper – at the question of whether the source supports the sentence that sits on top of it. Anyone publishing under their own name therefore needs a separate verification step. That applies to Perplexity, to our own pipeline, and to humans.
The Insight Took Three Minutes
First, the part that comparisons of this kind tend to leave out: on substance, both systems arrived at almost the same place. Independently, from the same prompt, both found the three-axis conflict (no meat wins on health, animal welfare, and environment at once), both found the poultry trap (switching from beef to chicken improves the climate footprint but multiplies the number of animals affected per quantity of meat – by a factor of roughly two hundred),4 both drew the line between processed and unprocessed meat, and both landed on “less meat” rather than “different meat” as the load-bearing frame.
More than that: the Perplexity report does three things better than our article. It treats fish as its own category, it mentions the nutritional detail of vitamin B12, and its side-by-side table of trade-offs is didactically strong.
So the core insight was available in three minutes. The reliability in the details was not – and that is no small thing, because the details are exactly where it becomes clear whether you can trust a text you are not going to re-research yourself.
What the Verification Found – Counted Honestly
We put the Perplexity report through the same evidence check as our own articles: every load-bearing claim is tested against a raw extract – a character-exact copy of the source’s raw text, pulled by a separate agent that does not know the claim under review.5 The most important result first: almost all the cited numbers really are in the sources. The sources exist, the links work. The report fails one layer deeper.
Counted honestly, five findings remain under our house rules – and this count is a convention, not an objective measure. Only two of them are hard, recomputable errors:
First, a unit error by a factor of 1,000. The report writes that a kilogram of beef requires “roughly 15,400 cubic meters of water” – that would be 15.4 million litres per kilo. The cited source, the Water Footprint Network, says 15,400 cubic metres per tonne, i.e. 15,400 litres per kilo.6 The kicker: the report’s own table lists the correct unit right next to it. Text and table contradict each other, and nobody saw it.
Second, a source does not support the sentence that sits on it. The report claims that the pooled data of six US long-term observation groups (cohorts) did not link poultry to cardiovascular disease. The cited study (Zhong et al. 2020, JAMA Internal Medicine) found the opposite: poultry consumption was significantly associated with cardiovascular disease – hazard ratio 1.04, meaning the disease rate was four percent higher, with an absolute 30-year difference of about one percentage point.7 Put precisely: the error lies in the rendering, not necessarily in the conclusion – the effect is small, and the study authors themselves were cautious because of possible preparation effects.8 But the report makes poultry the primary animal protein source in its recommendation – resting on a source that says the opposite of the rendered sentence, two paragraphs after its own description of the poultry trap.
Added to that are two findings that block under our house rules but are debatable: a suffering estimate (in essence: “broiler chickens spend the majority of their lives in pain”), presented as fact even though a 2025 review of the underlying methodology found only low overlap between independent raters and classified their estimates as “overprecise”;9 and cherry-picking from a UK Biobank study – the report quotes the most dramatic sub-figure (101 percent elevated stroke mortality risk) while omitting that the same study found no association with all-cause mortality.10 Finally, one point of source quality: a “10 to 100 times” claim rests on a LinkedIn post as its only support.
Across all findings, two patterns: not a single conflict of interest flagged – among the sources supporting meat nutrition claims is a poultry breeding company – and relative risks throughout, without absolute translation.
This matches what independent research says about the tool class: Perplexity Deep Research’s citation accuracy is high – an academic benchmark, i.e. a standardised comparison test, measures 90.24 percent correctly attributed citations.11 The weaknesses sit one layer deeper, in evidence coverage – the question of whether the source substantively supports the claim – and in one-sidedness, where an audit study of the tool class finds gaps across the board.12
The Look in the Mirror: Sixteen Errors of Our Own, Caught Internally
If your takeaway from reading this far is “so Perplexity just can’t do it”, you have read it wrong. Our own error list on the same topic is long enough.
While writing the meat article, our pipeline produced or found errors of class A at every verification stage – class A meaning, in our house rules: a false substantive claim, a source that does not support the sentence, an invented or uncovered number. Sixteen findings in total before the text reached a reader: four in the first quick audit, one in the word-meaning check, seven in the two-stage evidence check, three in the strict final audit, and one in its second round – a knock-on error from our own correction.3
The error classes are the same as in the Perplexity report. “Source does not support the sentence”: Perplexity wrote that poultry is healthy where the study found the opposite – we wrote “sudden cardiac death” into a statement about an EFSA opinion in which that term does not appear. “Number without supporting evidence”: the Perplexity report carries a 97.6-billion animal figure without citation or reference year – our draft carried billion-scale subsidy figures that appeared in none of our raw extracts. Both were removed.
The difference between the systems, then, is not who makes fewer errors. The difference is whose errors reach the reader.
The Punchline: The Review Commits the Error It Denounces
The most instructive moment of this comparison concerns us.
Our review of the Perplexity report originally listed six findings. Number six: an alleged internal contradiction, because the text says “more than 80 billion land animals slaughtered” in one place and “more than 97.6 billion” in another. Sounds like a contradiction – but isn’t one. 97.6 is greater than 80; both sentences can be true at the same time. And between the two values sits the estimate that Humane World for Animals attributes to the FAO: 92.2 billion land animals “kept and slaughtered” annually.13
The reviewer did not catch this. A red team did – a separate agent with the explicit brief to attack our own comparison thesis. The finding was downgraded; what remained justified is only that the 97.6 billion appears in its source without citation or reference year. Our review had thereby committed exactly the error it denounces: a dramatic diagnosis that does not survive verification of its own claim.
It was not the only incident of this kind. During the strict final audit of our own meat article, the reviewing agent’s fetch tool invented a time reference (“spring 2023”) that appears nowhere in the source. What stopped it was not a smarter model but the character-exact raw extract against which every claim has to run.3
The lesson from both incidents: verification is not an authority you set up once and then trust. It is a discipline – and it has to run against the verifier too. Its most reliable tool is the dumbest one: the verbatim raw text of the source, against which every model, every reviewer, and every author has to be measured.
What This Comparison Does Not Prove
If even the review errs – what does this comparison prove at all? Less than the headline promises. Four limits we would rather write down ourselves before a reader does.
First: this is a single case. One prompt, one run, one topic. It illustrates a pattern that independent research finds across the tool class – working links and high citation accuracy, but weak evidence coverage, one-sidedness, and a clear gap from expert standards.12 14 It does not prove it.
Second: our pipeline has correlated blind spots. It consists of Claude models throughout. Research shows that language models err together: a study of over 350 models found that models from the same provider or with the same base architecture make more strongly correlated errors – on one of the two leaderboard datasets examined (71 models), model pairs agreed on average 60 percent of the time when both were wrong – and that, of all things, more accurate models correlate even more strongly, including across provider boundaries.15 On top of that, AI judges evaluate their own outputs with a bias: some systematically favour them, others penalise them.16 “Caught internally” therefore does not prove “nothing missed”. Our countermeasure is the model-agnostic raw extracts – they mitigate the problem, they do not remove it.
A side note on this: according to Perplexity’s help centre, Deep Research uses the models “Opus 4.6 Thinking” (Max subscription) and, respectively, “4.5 Thinking” (Pro subscription)17 – model names from the Claude family, as the trade press reports, which dates the switch to early February 2026.18 Our test run was on a Pro subscription; we did not document which model it actually hit. So quite possibly two Claude-based systems competed against each other here – which sharpens the correlation question rather than weakening it.
Third: the comparison was not blind. Our reviewers had a week’s head start in topic knowledge from working on our own meat article, complete with 21 finished raw extracts and a list of known problem numbers. At least one finding – the suffering estimate – was findable only because of that. So the comparison partly measures a domain head start, not pipeline superiority.
Fourth: there is a genuine counter-example. In literature search, human-compiled reference lists are clearly outperformed: only 51 percent of human citations were judged moderately relevant or higher, against 86 to 88 percent for the strongest AI-based re-rankers; at the same time, human authors are 2.5 times more likely to cite their own co-authors.19 Human source selection – ours included – is no reliable yardstick either, no “ground truth”.
Which leaves the question a decision-maker actually asks: does the effort pay off anyway?
The Bill: When Which Path Is Worth It
The cost asymmetry is substantial – and it belongs on the table. Perplexity Pro costs 20 US dollars a month – converted, about 17.50 euros; we could not verify an official euro list price.20 The report was done in minutes. Our approach cost about one working day and roughly 1.7 million agent tokens; the bill for the day it was created came to 188.49 US dollars, converted about 162 euros – of which roughly a third to a half went to the article, the rest to other work.3
By the metric “errors found per euro and minute”, Perplexity wins this comparison clearly. That is stated here in so many words.
Our thesis demands a different metric. For texts that appear under your own name or carry decisions, what counts is not the price per error found, but whether an error appears with your name above it. A factor-of-1,000 error in an internal research result costs a correction. The same error in a published text costs credibility – the currency in which consultants, expert authors, and decision-makers are paid.
This yields a simple division of labour: Deep Research for exploring, for internal orientation, as a starting point – it excels at that, and it is cheaper by orders of magnitude. And a separate verification step for everything that gets published – regardless of whether the text comes from Perplexity, from your own pipeline, or from a human.
The Market Is Pricing Depth Right Now
That depth costs compute is not a claim of ours. The vendor itself has been charging for depth since this year – and the development can be dated step by step:
At launch on 14 February 2025, Perplexity advertised Deep Research as free for everyone, with “unlimited Deep Research queries” for Pro subscribers.1 In early February 2026, the Pro quota was cut to about 20 runs per month – the starting figure sits, depending on the source, at around 500 to 600 runs per day; the cut itself is corroborated by multiple independent sources and hit existing customers mid-billing-cycle.21 The official reasoning in the help centre: “Usage limits have been adjusted to allocate more computing power per session, so you get deeper insights and higher-quality results.”17 Whether quality per run actually rose is open – no independent before-and-after measurement exists.
In parallel, the research yardstick is shifting: the newer benchmarks increasingly separate “does the link work” from “does the source support the claim” – exactly the layer this article is about.11 14
Both point to the same conclusion: the market is moving from “plenty and fast” to “less, but deeper” – and the evaluation standards are migrating from the citation surface into the evidence depth. Anyone building a publication pipeline today should take that movement seriously: the research system delivers the raw material. Verification against the source’s raw text – separate, repeatable, and run against itself too – turns it into something you can sign.
Condensed into three questions before a deep-research report goes out the door:
- Does every load-bearing source support the sentence that sits on it – checked against the source’s raw text, not against a working link?
- Are the sources’ conflicts of interest flagged, and does every relative risk come with the absolute number next to it?
- Did the verification also run against the verifier – with a research brief set against your own thesis?
Anyone who answers yes to all three may sign. Because citing is not verifying – the second step is the one you stake your name on.
To close, the transparency offer we would demand of any third-party report: the unabridged Perplexity report is published as an evidence document (linked at the top of this article); all source raw extracts and our complete verification log – including the finding our own review did not survive – are on file and will be disclosed on request. The same falsifiability this article demands.
Quellen
-
Perplexity: “Introducing Perplexity Deep Research”, blog post of 14 February 2025. Original URL https://www.perplexity.ai/hub/blog/introducing-perplexity-deep-research (blocks automated access); checked verbatim against the Wayback snapshot from launch day: https://web.archive.org/web/20250214203504/https://www.perplexity.ai/hub/blog/introducing-perplexity-deep-research (accessed 30 August 2026). Verbatim quotes: “expert-level analysis across a range of complex subject matters”, “completing most research tasks in under 3 minutes”, “Pro subscribers get unlimited Deep Research queries”. Vendor self-description, and therefore interest-driven. A side note on the diligence of secondary sources: the launch post states 20.5 percent for the “Humanity’s Last Exam” benchmark; the tech press and Perplexity’s own X channel circulated 21.1 percent – even the self-description exists in two versions. ↩ ↩2 ↩3
-
“Not Which Meat, but How Much: Why the Best Compromise Is a Quantity Target – Not a Label” (A16), this blog. The article our pipeline produced from the same prompt. ↩
-
Internal run logs (harness logs) and verification logs of the article’s creation (agent runs, token billing, finding lists of all six verification stages, day’s bill of 29 August 2026). Disclosed on request; euro conversion at the ECB reference rate of 28 August 2026 (0.85889 EUR/USD). ↩ ↩2 ↩3 ↩4
-
Bryant Research: around 200 chickens replace the meat quantity of one cow; checked against the source in the A16 article. See the footnote there, including the source reference. ↩
-
Methodology reference: two-stage evidence check – one agent extracts the source blind (does not know the claim under review), a second compares claim and extract without net access. Prompt templates in the open repository: https://github.com/flodido/agentic-research-vault ↩
-
Water Footprint Network: “What can consumers do?”, https://www.waterfootprint.org/time-for-action/what-can-consumers-do/ (raw extract of 30 August 2026). There: 15,400 m³/t for beef. The report under review writes “roughly 15,400 cubic meters of water” per kilogram – a thousand times as much. ↩
-
Zhong, V. W. et al.: “Associations of Processed Meat, Unprocessed Red Meat, Poultry, or Fish Intake With Incident Cardiovascular Disease and All-Cause Mortality”, JAMA Internal Medicine, 3 February 2020. https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2759737 (raw extract of 30 August 2026). Poultry: HR 1.04 (95% CI 1.01–1.06) for incident cardiovascular disease; no significant association with all-cause mortality. ↩
-
Northwestern Feinberg School of Medicine (press office of the study institution): “Meat Consumption Raises Risk of Heart Disease and Death”, 3 February 2020. https://news.feinberg.northwestern.edu/2020/02/03/meat-consumption-raises-risk-of-heart-disease-and-death/ (accessed 30 August 2026) – confirms the poultry finding and the authors’ caution (possible preparation effects, chicken skin). ↩
-
McAuliffe, W.: replication attempt of the Cumulative Pain estimates, Rethink Priorities, 3 November 2025. https://rethinkpriorities.org/research-area/cumulative-pain-framework/ (accessed 29 August 2026). Three independent experts replicated four estimates; the overlap of the estimate intervals was low (“Overlap in average time in pain was fairly low, suggesting that raters’ raw estimates were overprecise”). The underlying methodology comes from the Welfare Footprint Institute (advocacy-adjacent research). ↩
-
UK Biobank: “Red meat consumption and all-cause and cardiovascular mortality: results from the UK Biobank study”. https://www.ukbiobank.ac.uk/publications/red-meat-consumption-and-all-cause-and-cardiovascular-mortality-results-from-the-uk-biobank-study/ (raw extract of 30 August 2026). The study found elevated cardiovascular sub-risks, “but not all-cause mortality”. ↩
-
DeepResearch Bench: “A Comprehensive Benchmark for Deep Research Agents” (arXiv:2506.11763), 2025. https://deepresearch-bench.github.io/ – FACT framework: 90.24 percent citation accuracy for Perplexity Deep Research; other Perplexity modes far below (Sonar Reasoning Pro: 39.36 percent). Independent academic source, methodology disclosed (100 PhD-level tasks, 22 fields). ↩ ↩2
-
Venkit, Laban et al. (Salesforce AI Research): “DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence”. https://arxiv.org/abs/2509.04499 (accessed 30 August 2026). Finding across several systems: citation accuracy 40–80 percent, high one-sidedness scores. Note the vendor proximity: Salesforce runs an agent business of its own. ↩ ↩2
-
Humane World for Animals: “More animals than ever before – 92.2 billion – are used and killed each year for food”. https://www.humaneworld.org/en/blog/92-billion-animals-used-killed-meat-each-year (accessed 30 August 2026). NGO source (advocacy); used here solely as evidence of an estimate sitting between the two values (“An estimated 92.2 billion land animals are kept and slaughtered annually in the global food system, according to the Food and Agriculture Organization” – no reference year). ↩
-
“DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation” (arXiv:2512.17776) and DeepResearch Bench II (arXiv:2601.08536, both accessed 30 August 2026): even the strongest deep-research agents remain clearly behind experts in expert grading rubrics. Complementary on the tool class: “Cited but Not Verified” (arXiv:2605.06635) separates working links from actual evidence coverage; Perplexity itself is not tested there. ↩ ↩2
-
Kim, Garg, Peng, Garg: “Correlated Errors in Large Language Models”, ICML 2025. https://arxiv.org/abs/2506.07962 (accessed 30 August 2026) – analysis of over 350 models across two leaderboards and a résumé-screening task. On the Helm dataset (71 models), wrong answers of model pairs agree on average 60 percent of the time (chance expectation: one third). Models from the same provider or with the same base architecture correlate more strongly; individually more accurate models correlate even more strongly, “even with distinct architectures and providers”. ↩
-
“Quantifying and Mitigating Self-Preference Bias of LLM Judges”. https://arxiv.org/abs/2604.22891 – measures the self-preference bias of AI judges towards their own outputs: widespread, but not uniformly self-favouring (of 20 models, 8 biased positively, 9 negatively, 3 neutral); family effects across model boundaries are explicitly named by the paper as an open, unresolved limitation. ↩
-
Perplexity Help Center: “What’s New in Advanced Deep Research”. Original URL https://www.perplexity.ai/help-center/en/articles/13600190-what-s-new-in-advanced-deep-research (blocks automated access); checked verbatim against the Wayback snapshot of 11 February 2026: https://web.archive.org/web/20260211012137/https://www.perplexity.ai/help-center/en/articles/13600190-what-s-new-in-advanced-deep-research (accessed 30 August 2026). Verbatim: “Max subscribers get Opus 4.6 Thinking”, “Pro subscribers get 4.5 Thinking gradually”, “Usage limits have been adjusted to allocate more computing power per session, so you get deeper insights and higher-quality results.” The page carries no absolute modification date. ↩ ↩2
-
Dataconomy: “Perplexity Upgrades Deep Research Tool With Claude Opus 4.5 Integration”, 5 February 2026. https://dataconomy.com/2026/02/05/perplexity-upgrades-deep-research-tool-with-claude-opus-4-5-integration/ (accessed 30 August 2026) – attributes the help centre’s model names to Anthropic’s Claude family. ↩
-
Sahu, Charlin, Pal: “Rethinking Literature Search Evaluation: Deep Research Helps, and Human Citation Lists Are Not a Ground Truth”. https://arxiv.org/pdf/2605.29234 (accessed 30 August 2026). ↩
-
GIGA.de: “Perplexity-Abos im Check”, 6 January 2026. https://www.giga.de/tech/perplexity-abos-im-check-wann-sich-das-upgrade-wirklich-lohnt—01JYE759P51Z7MWWKCRSEA55VF (accessed 30 August 2026). The euro figure (~€17.50/month) is recognisably a conversion of the US price of $20, not a separate EU list price; VAT treatment unclear. The article is marked as advertising (“Anzeige”) and carries an affiliate link disclosure by the editorial team. ↩
-
MakeUseOf: “If you bought an annual Perplexity subscription, you were lied to”, February 2026. https://www.makeuseof.com/bought-annual-perplexity-subscription-lied/ – independently corroborated by the limit tracker verified against the Wayback Machine, https://ailimit.watch/tools/perplexity/ (as of 6 June 2026, accessed 30 August 2026: “Deep Research quota slashed ~500/day → 20/mo”, flagged as unannounced). The starting figure sits, depending on the source, at around 500 to 600 per day (MakeUseOf: “500 to 600 Deep Research queries per day”); the cut to ~20/month and the lack of advance notice to existing customers are corroborated by multiple independent sources. ↩
Transparency notice: This article was created with AI assistance and reviewed editorially before publication.