PASSIONFRUIT PULSE
The "AI visibility score" era is ending faster than the dashboards selling it.
This week we published research synthesizing 25+ major studies, 680M+ tracked citations, and 2.2M+ ChatGPT responses. The conclusion underneath all of it: the metric the entire AI tracking industry is built on — citation count — is noisy, often misleading, and missing the variables that actually predict revenue.
The same week, Bing rebuilt its AI Visibility Reports around Intent, Topics, and Citation Share rather than one score. Our own blog opened a layer most teams haven't started thinking about making your site usable by AI agents, not just visible to them.
Three signals, one direction: the work is splitting into layers. Teams measuring it as one thing are about to lose ground to teams measuring it as four.
Here's what we're tracking.
Blog of the week
The next layer most teams haven't started: Making your site agent-ready

Citation work earns you a mention. Agent-readiness lets an autonomous agent actually use your website to complete a task book, compare, fill, purchase without a human ever loading the page.
Google's developer guidance made the shift official this week. The web.dev article "Build agent-friendly websites" tells developers they now have "a new type of visitor." MCP (Anthropic's Model Context Protocol) hit 110M monthly downloads — faster than React reached the same milestone. Chrome's WebMCP entered origin trial in May, backed by Google, Microsoft, Booking.com, and Shopify.
The strategic detail most marketers miss: agent-readiness is largely the WCAG accessibility playbook restated for a second audience. Google's own line: "Everything we suggest making a site agent-ready also makes sites better for humans."
The seven concrete rules from Google's checklist all map to existing accessibility practice: semantic HTML, stable layouts, cursor:pointer signals, properly linked form labels, minimum 24×24 interactive targets, no ghost overlays, visible state changes on action. Brands maintaining accessibility have done most of the work. Brands that haven't just got a second business reason to do it.
The protocol stack MCP, WebMCP, UCP (Google's commerce protocol), AP2 (Google's payment protocol) — is early and unsettled. The right move is not to bet on a winner. It's to ship the accessibility-tree foundation that pays off across all of them.
The two layers stack. Citation work earns the recommendation. Agent-readiness lets the agent act on it. Brands doing only citation risk being recommended but unable to convert. Brands doing only agent-readiness risk being usable but never recommended.
Research you should read
The research: The AI tracking industry doesn't want you to read

The headline finding, drawn from SparkToro's 600-volunteer, 2,961-prompt test: less than 1 in 100 chance the same brand list appears twice on ChatGPT. Less than 1 in 1,000 chance it's in the same order. "Ranking position in AI" is, statistically, noise. Rand Fishkin's blunt line in the paper: any tool that gives you one is full of baloney.
It gets worse before it gets better. Five other findings worth internalizing:
Ghost citations are everywhere. Seer Interactive analyzed 541,213 LLM responses. One client's blog was cited 100+ times in 25 days with zero brand mentions. Every citation was a footnote supporting the competitor's recommendation. Citations are the bibliography, not the brainstorm.
85% of content AI retrieves never shows up in the answer. AirOps analyzed 548,534 retrieved pages across 15,000 prompts. Only 15% became citations. Discovery isn't visibility.
Citations themselves are frequently wrong. Nature Communications: 50–90% of LLM responses weren't fully supported by the sources they cited across 58,000 statement-source pairs. The model presents 90%+ confidently anyway.
Cross-platform overlap is 11%. A single-platform score is a partial signal from a system that disagrees with every other system 89% of the time. Superlines found the same brand can see 615× citation volume difference between Grok and Claude.
Citations are temporally unstable. ChatGPT cited Reddit in nearly 60% of responses in August 2025; by mid-September that collapsed to 10%. One platform change wiped out years of Reddit optimization.
The framework that works is a measurement stack, not a number:
Recommendations (brand actively suggested) → parametric memory + brand. Owned by CMO. The signal that compounds.
Mentions (brand named, no link) → training data + retrieval. Owned by PR.
Citations (URL linked) → real-time retrieval. Owned by content. Volatile.
Used vs. cited split (informing the answer invisibly vs. attributed) → owned by GEO.
Citations are the most-tracked, least-stable, least-meaningful layer. Recommendations are the inverse. If your dashboard reports one number, you're measuring the most volatile variable as if it were the most important one.
INDUSTRY SIGNAL
Industry signal: Google's first AI Mode usage report is a content brief in disguise

Chris Long at Nectiv flagged this week that Bing's AI Visibility Reports just added four new dimensions:
Intent classification per grounding query (Informational, Commercial, Navigational, Research)
Topics via AI/ML classification (least useful per Chris — most bucket as "Marketing" or "Business")
Citation Share — for each grounding query, the percentage of times you get cited
Compare — visibility deltas over time
The most operationally valuable is Citation Share. Instead of "what's my AI ranking," you now see whether you reliably surface for specific category questions — and how that trends. This is exactly the visibility percentage SparkToro's research argued for.
The signal underneath the launch matters more than the feature itself. A major search platform is rebuilding its AI reporting around multi-dimensional signals at exactly the moment the research is converging on the same answer. Single-score AI dashboards aren't going to age well. Bing leading here means the rest will follow.
TOOLS
Product of the Week: How do you turn this into work that actually ships?

The hard part of this new measurement stack isn't generating the data. It's converting it into prioritized work that ships across content, technical, research, and PR every week, on the right pages, in the right order.
The view we run internally: Auto Task Suggestions. A workflow surface that turns every AI Visibility gap and Traffic Drop signal into a specific, prioritized, accept-or-reject task — mapped to a page, with the category attached.
A snapshot from a recent client workspace targeting monday.com competitive territory: 237 tasks generated, 55 accepted, 102 rejected, 80 pending. Sample tasks at the top of the priority list:
"Add monday.com AI Features section covering Sidekick, AI Blocks, and AI Workflows" (On Page, Traffic Drop)
"Noindex the White Paper PDF — splitting ranking signals from the main scheduling page" (Technical, Traffic Drop)
"RESEARCH: Confirm monday.com's AI Work Platform repositioning has displaced [legacy positioning]" (Research, Traffic Drop)
"Submit a data-led expert quote to projectmanager.com for inclusion in their Primavera comparison" (PR Outreach, AI Visibility)
Notice what this is not: an AI visibility score moving up or down. It's the work layer the score should drive. Each task is tagged by source, category, page, priority. The team accepts, rejects, or assigns. Nothing sits in a dashboard waiting to be interpreted.
This is the gap most AI visibility tools never close. Showing you a number is easy. Telling you which page, which section, which schema fix, and which PR pitch to send next is the work.
The Passionfruit Playbook
One question to audit: When your AI tracking vendor reports a "ranking" or single visibility score, ask: across how many runs per prompt? If the answer is under 60, you're paying for noise.
One quick win: Run Lighthouse 13.3's Agentic Browsing audit on your top revenue page. Fix the highest-impact accessibility failures. Same fix improves accessibility, SEO, and agent-readability simultaneously.
One thing to stop doing: Reporting AI visibility as one number. Replace it with the four-layer stack: recommendations, mentions, citations, used-vs-cited. The complexity is the point.
One creator prompt: What's one piece of content I publish this week that's structured to be recommended (parametric, brand-name-bearing) rather than just cited (referenced in passing)? The two are different jobs.
Until next week,
Passionfruit Team



