PASSIONFRUIT PULSE

This issue's lead piece is the cleanest frame we've published for understanding what's happening to content visibility in 2026: The Citation Paradox.

The thesis in one line: generation and attribution are opposing operations. A language model produces text by moving toward the high-probability center of its training distribution. A retrieval system attributes a source in proportion to the information that source carries beyond what the model could already generate. The same content cannot be optimized for both.

That's not a quality argument. It's a redundancy argument — and it explains, with peer-reviewed math underneath it, why your team can publish more, polish harder, ship cleaner schema, and watch AI visibility stay flat.

Here's what we're tracking.

Blog of the week
Google Ads just stopped showing you what users actually typed

Google quietly updated its Ads documentation to clarify that for searches through AI Mode, AI Overviews, Google Lens, and autocomplete, the term shown in your Search Terms Report is now Google's "best approximation of the user's intent" — not the literal query.

In plain terms: for a growing slice of search activity, the term in your report is Google's read on what the user wanted, not the words the user entered. The change isn't a feature launch. It's Google formalizing behavior that's been emerging for a year.

What breaks:

  • Negative keywords: adding an interpreted term may not block what you think it blocks. Reported text and actual query are no longer guaranteed to match.

  • Compliance review: for legal, healthcare, and finance accounts, the gap between excluded terms and inferred intent is real brand-safety exposure.

  • Customer language mining: exact phrasing gets smoothed out. Qualitative insight is lost.

  • Internal reporting: presenting reported terms as "customer language" to execs is no longer accurate without a caveat.

The pattern underneath the announcement matters more than the announcement itself. Journey-Aware Bidding moved bidding from single events to journeys. AI Mode ads moved creative from static assets to Gemini-generated ones. Now reporting shifts from literal queries to interpreted intent. Every move points the same direction: Google Ads is shifting from advertiser control toward AI interpretation of intent.

Accounts on broad match plus Smart Bidding barely feel it. Accounts built on exact-match keywords and granular negative lists will feel friction across all of these. The path forward for paid and organic both: strong first-party tracking and independent measurement that doesn't depend on Google interpreting things for you.

Research you should read
The Citation Paradox: why "good enough" AI content fails the citation test

The argument bridges three research literatures that have been developing in isolation:

  • Recursive model collapse (Shumailov et al., Nature 2024): models trained on their own outputs degrade across generations as low-probability tails disappear.

  • Alignment-induced diversity loss (Kirk et al., ICLR 2024): RLHF measurably narrows the band of text aligned models produce.

  • Co-writing homogenization (Padmakumar & He, ICLR 2024): independent authors using the same aligned model produce essays measurably more similar to each other.

All three say the same thing: generated text whether produced by a model, an aligned model, or a human using one converges toward a center.

Now the other side of the paradox. The most rigorous GEO study to date (Aggarwal et al., KDD 2024) found that the modifications most likely to earn AI citations are specific statistics, direct quotations, and citations to authoritative sources — exactly the features that mark content the model couldn't have produced on its own.

The strategic consequence is harsh and clean:

The probability that a generative engine attributes a passage rises with the passage's surprisal relative to the engine's prior. The surprisal of generated content is, by construction, low. Producing content by generation lowers, in expectation, the very quantity that citation rewards.

The under-discussed implication: in the pre-AI web, producing competent average content was expensive and that expense was a moat. Generative models collapsed the marginal cost of average content to roughly zero. Average content stopped being scarce, which means it stopped being a moat. The durable advantage shifts to whatever still has positive marginal cost and cannot be generated: original data, first-hand testing, proprietary results, named expertise, defensible positions.

The brands that win AI citation produce the things models are trained on, not the things models produce.

INDUSTRY SIGNAL
Industry signal: Google's first AI Mode usage report is a content brief in disguise

Worth pairing this week's research with the official Google Search Central announcement from earlier this quarter their guide to optimizing for generative AI search.

The list of tactics Google itself called unnecessary llms.txt files, AI-specific content chunking, AI-specific rewriting, manufactured citations, special AI schema is interesting on its own. The list of what Google says actually works is more revealing. It maps almost perfectly to what the Citation Paradox predicts:

  • Non-commodity content with a real point of view = high-surprisal information the model can't generate.

  • Earned editorial coverage and community participation = testimony from authorities the model defers to.

  • Multimedia and entity signals = named, first-hand evidence that carries provenance.

The Citation Paradox gives this list its theoretical basis. Google says, in effect, "publish things models can't fake." The Paradox explains why: models attribute sources that supply information beyond what they can vouch for from training. Aligned-model output cannot occupy the terminus of that deference chain — it has no provenance other than the corpus it averages.

Two independent paths Google's product docs and peer-reviewed information theory pointing at the same operational answer. That's signal.

TOOLS
Product of the Week: Are you even in the AI conversation?

Most teams measure whether they're cited. The more useful measurement is what specifically gets pulled into the answer and at what mention rank, with what sentiment.

A snapshot from a recent Passionfruit Labs workspace tracking mentions across May 2026:

  • May 3 (ChatGPT, rank 1, +20 sentiment): cited for "in-house vs outsourced SEO cost figures"

  • May 6 (ChatGPT, rank 1, +20 sentiment): cited as "authoritative guide on 2026 SEO pricing and in-house vs agency costs"

  • May 6 (ChatGPT, rank 3, +40 sentiment): SEO.com cited for "questions to ask a prospective SEO agency, including pricing and deliverables"

  • May 27 (ChatGPT, rank 1): cited for "cost ranges and hybrid strategy guidance in the in-house vs agency comparison"

Read the contexts carefully. The citations aren't for general advice. They're for specific numbers and frameworks — cost figures, pricing ranges, comparison structures, hybrid strategy guidance. Exactly the surprisal-rich inputs the Citation Paradox predicts get cited. The brand isn't getting picked up for "what is an SEO agency." It's getting picked up for things only its own data and operator experience could produce.

This is the measurement that matters: which specific claims, with what proprietary information underneath them, are being attributed. Track that, and you know which content assets are doing the citation work and which get retrieved but never credited.

The Passionfruit Playbook

One question to audit: Of the last 10 things AI engines cited from your domain, how many were because of proprietary information specific numbers, named examples, original analysis versus general topic coverage? If you don't know, you're optimizing the wrong half.

One quick win: Take your top-traffic blog post. Add one specific, proprietary number this week — a result from your own data, a number from a client engagement, a finding from a test you ran. That single surprisal injection beats a full month of polish.

One thing to stop doing: Iterating on quality to fix a citation problem. The deficit isn't quality. It's redundancy. Better generation makes redundancy worse.

Creator prompt: "What's the one piece of information I have access to that no model in the world has seen yet? Publishing that beats a hundred polished explainers."

Until next week,

Passionfruit Team

Keep reading