Cited does not mean quoted: an audit of 98,020 AI Overview claims
marketing October 10, 2026 · Mintec

Cited does not mean quoted: an audit of 98,020 AI Overview claims

WashU audited 7,583 AI Overviews against the pages Google cited: 11% of claims unsupported, only 41.9% fully grounded. What claim fidelity means for GEO.

Cited does not mean quoted: an audit of 98,020 AI Overview claims

Mostly, but not always. An independent audit of 7,583 AI Overviews found that about 89% of their claims were supported by the pages Google cited, 11% were not, and only 41.9% of overviews had every claim supported. A citation tells you your page was used. It says nothing about whether the sentence sitting next to your link is something you actually wrote.

That gap has a name in the paper: claim fidelity. It is the difference between being in the source list and being represented correctly, and until this study nobody had measured it at scale. The work comes from Haofei Xu, Umar Iqbal and Jacob Montgomery at Washington University in St. Louis, and it is being presented this week at the ACM Internet Measurement Conference in Karlsruhe.

What they actually measured

Over 40 days, from March 13 to April 21, 2026, the team ran 55,393 trending queries drawn from Google Trends across 19 categories, from servers in Northern Virginia with a fresh cookie-less browser each time. They captured 7,583 AI Overviews, the reference citations inside each one, the first-page results shown alongside, and the full text of every cited page.

Then the part nobody had done: they cut each overview into one-fact-per-sentence claims, rewrote each to stand on its own, and checked all 98,020 of them against the cited pages.

Two details about the verification matter. A model (Grok 4.1 Fast Reasoning) made the judgements, and it agreed with human annotators on 98 of 100 double-checked claims. And the rewriting step is the interesting one: the researchers had to replace pronouns with full entity names and drop section headings to make claims checkable. More on that below.

The number people will quote is the wrong one

The study found AI Overviews on 13.7% of queries. That number is already losing a fight with the 60% figure that tool vendors have been circulating since November 2025, which comes from a keyword tracking set of a few thousand terms rather than 55,393 trending queries. Same feature, two completely different denominators.

Activation also depends on how you phrase the question. Question-form queries triggered an overview 65% of the time against 9.5% for everything else. Among non-question searches the rate climbed from 9.9% for one-word queries to 38.7% for six words or more. Hobby and leisure searches got one 46.1% of the time; political searches, 7.5%.

If you remember only one thing from the methodology: any AI visibility percentage you are shown needs its denominator attached, or it is decoration. We make that argument at length in why most AI visibility scores are misleading.

Credible sources, inaccurate sentences

The source quality findings are genuinely good news for publishers. Across 61,212 citations, the domains Google pulled into overviews scored higher on average credibility than the first-page results shown for the same searches, and leaned less on user-generated content. The average overview cited about eight sources.

Accuracy is where it falls apart. Of the 98,020 claims:

  • 84.6% were clearly supported by the cited text
  • 4.4% were vaguely supported
  • 7.0% could not be found in the cited page content the researchers retrieved
  • 2.7% contradicted the cited source
  • 1.4% cited a page that contradicted itself

Only 41.9% of overviews were fully grounded, meaning every verifiable claim was supported. Two caveats keep this honest. The 11% is an upper bound: the pipeline could not collect Reddit threads, YouTube videos, Facebook or Instagram posts, and the authors estimate the rate would drop to about 5.3% under the most generous assumption. Fast-moving facts like weather or school closures may also have changed between the overview and the page capture.

Still, roughly six in ten overviews carry at least one claim their own sources do not back. That is the sentence I would put on a slide.

One more finding deserves more attention than it gets. Only 25.0% of an overview's reference domains overlap with the top five organic results, 41.4% with the top ten, and 70.2% with the full first page. Overall, 29.8% of cited domains did not appear on the first page at all.

That is a second, separate source-selection system running beside the ranking you already spend your week on. It explains the persistent complaint that pages with modest rankings get cited while page-one incumbents do not. It is also why the three-engines problem we described in AI search runs on three citation rulebooks does not resolve by ranking first in Google.

The part that should change how you write

Here is the reversal. The auditors had to rewrite your sentences before they could verify them. The model reading your page is doing the same work, with worse judgment: it is reconstructing what you meant from context, then handing a paraphrase to the reader with your domain attached.

So write for the lift. Concretely, five checks:

  1. One fact per sentence, with the entity named. If your claim needs the previous sentence to make sense, it will not survive being extracted.
  2. The number travels with its denominator and its date. "Conversion rose 40%" is a claim waiting to be misquoted. "Conversion rose 40% among 1,200 trials in Q2" is not.
  3. Scope the claim to the data. A global statement resting on a US sample is the most common unsupported claim we see in audits.
  4. Keep the source in the same paragraph as the number. Citations stranded three hundred words away get dropped when the paragraph is extracted.
  5. Read the sentence you would want quoted. If you would not sign it, do not publish it.

This is not a new philosophy, it is the same standard Google itself put in writing on October 1 when it called manual fact-checking of AI content "critical".

Where you can see it on your own site

Search Console's Generative AI report gives impressions by page, country, device and date. It gives no queries, no clicks, and no way to see the paraphrase. So the check has to be manual:

  1. Take your ten pages with the most impressions in that report.
  2. Ask their target query in AI Mode and expand the overview.
  3. Compare the sentence attached to your link with your page.
  4. Log the three failure modes: missing denominator, over-broad scope, conclusion you never wrote.

Twenty minutes, once a week. It is the only way to catch a misquote before a customer or a competitor does. For the measurement layer around it, see how to measure GEO performance.

What the study does not prove

The paper is marked under review and has not passed peer review. The query set is trending US searches from a 40-day window in early 2026, not your customers' actual questions. Verification was automated, however carefully calibrated. And the publisher revenue concern in the paper is a hypothesis: the team found that 2.2% of pages with an overview also showed a Google sponsored ad, with 39 cases of an ad sitting above the overview, but did not measure lost traffic or lost money.

None of that changes the core result, because the claim-fidelity measurement does not depend on the traffic question at all.

What to do this week

Run the ten-page check above and write down what you find. Then fix the writing, not the tracking: most unsupported claims are not model errors, they are sentences that were allowed to be vague. The teams worth worrying about are the ones whose pages are cited and whose claims are clean, because they are about to be quoted correctly everywhere while yours is not. For the framework this sits inside, see SEO, AEO and GEO as one system.

Frequently Asked Questions

What share of AI Overview claims are not supported by the cited page?

About 11% of them, in the largest independent audit so far. Researchers at Washington University in St. Louis split 7,583 AI Overviews into 98,020 single-fact claims and checked each one against the pages Google cited. Roughly 89% were supported, 7% could not be found in the retrieved text of the cited page, 2.7% contradicted it, and 1.4% cited a page that contradicted itself. The authors call 11% an upper bound, because they could not collect content behind logins or social platforms; under their most generous assumption the figure falls to about 5.3%.

Does an AI Overview citing your page mean it represents you accurately?

No. A citation means the page was retrieved and used, not that the sentence next to the link matches what the page says. In the same audit only 41.9% of AI Overviews were fully grounded, meaning every verifiable claim was supported by the available cited text. The practical exposure is that a reader, or a competitor, can now see a paraphrase of your claim with your domain attached to it, and you have no editorial control over that sentence.

How do you check whether an AI Overview is misquoting your content?

Open Search Console, go to Performance, then Generative AI features, and take the pages with the most impressions. Run their target query in AI Mode, expand the overview, and compare the sentence attached to your link with what your page actually says. Check for three failure modes: a number without its denominator, a claim scoped more broadly than your data, and a conclusion your page never states. Ten pages takes about twenty minutes.

Related Articles