The quarterly traffic review shows direct visits flat or declining, and no one can pinpoint the cause. The dashboard shows the dip and hides the reality: the buying decision now happens in an AI assistant’s chat window, before a customer ever types your brand, or a competitor’s, into a browser. Your analytics has no row for that conversation. And some of the tools sold to show it to you never ask the AI what it is saying.
I argued in AI Is an Audience, Not a Channel that AI assistants are an audience your brand has never measured. That piece asked what the assistants should say about your brand. This one looks at where they are recommending someone else, the invisible cost of those lost recommendations, and who inside your organization has to own the fix.
Where the decision moved
Immad Akhund, co-founder and CEO of Mercury, described the shift in August, talking about money on his own podcast: “I don’t think the LLMs need to… do the math or… do the money movement or whatever, but they can direct money to the right place for people.” (LLMs are the large language models behind assistants like ChatGPT.)
Directing money to the right place is the entire shift. The assistant does not need to close the sale. It only needs to decide which product the customer sees first, which alternatives appear alongside it, and the reason it gives for the recommendation.
Buyers arriving from these assistants behave differently. Adobe measured AI-referred traffic to U.S. retail sites in May: a 138% increase from a year earlier, and a purchase conversion rate 54% better than traffic from other sources. A buyer arriving from an AI assistant has already been told what to buy and where to buy it. If the assistant recommends a competitor, that buyer lands on the competitor’s site instead, with the decision made. If you sell anything a customer compares before buying, this conversation is happening in your category, whether or not the numbers have reached your dashboard yet. If your AI-referred traffic is flat while it grows this fast across U.S. retail, the conversation is happening and a competitor is winning it.
Customers diverted by an assistant never reach your site. Your funnel records no abandoned carts or bounced visits, because it never knew these buyers existed. In early 2025, Bain & Company found that roughly 60% of searches already ended without a click to any website. For those people, the conversation was the visit. Every customer an assistant routes to a competitor or a retailer is a sale you never knew you lost.
The analytics blind spot
Web analytics counts arrivals based on referrers, the record of which page sent a visitor to you. A decision made inside an AI chat has no referrer. When an assistant names a competitor as the safer choice, links to a marketplace rather than your own store, or describes your brand with a positioning you retired years ago, nothing is logged in your systems.
The damage compounds. A wrong answer costs you the sale, then lands on your customer service team as a claim your company never made, with no way to trace where it came from.
Your instruments also cannot see how an assistant builds its answer. It draws on two sources: what it absorbed about your brand during training, which I call the memory read, and what it finds through web search at the moment of the prompt, the live read. These are two separate problems, with different fixes on different timescales. Two more complications: the free and paid versions of the same assistant can disagree about you, and competing assistants disagree with each other. There is no single “AI answer” to check.
The reality is a grid: several assistants, free and paid tiers, and separate memory and live reads, across scores of distinct customer prompts. One screenshot from one prompt on one day is an ‘n=1’ sample. It is not measurement.
What the measurement tools miss
The market for answer engine optimization (AEO) software is expanding rapidly. AEO is the work of changing how AI assistants describe and recommend you; generative engine optimization (GEO) is the same work under another name. Ara Kharazian, lead economist at Ramp, said in May: “AEO is the software that firms use to track their performance in AI models and whether or not they’re being recommended. Huge growth.”
I compete in this category, so read the next few paragraphs with that in mind.
Investors agree. On Tuesday, Profound announced a $180 million funding round at a $1.8 billion valuation, less than seven months after its previous round, and said its revenue had tripled in six months. The market has decided that tracking AI mentions of a brand is worth paying for. Profound’s own product page shows what that tracking is: an industrial-scale mention dashboard that captures visibility, share of voice, and web citations directly from the browser, refreshed daily.
That has real value, and it answers the wrong question. Because it captures answers only with web search switched on, it never separates the memory read from the live read, and those are two different problems with two different fixes. It reports sentiment, but it cannot tell you whether the assistant associates your brand with the positioning you intend. And a dashboard does not diagnose the problem, assign the fix, or put a person’s name under a recommendation.
More concerning are the tools that bypass the assistants entirely. Instead of querying the models, they run clickstream data, the records of pages that large panels of internet users visited, through a statistical model to estimate what an assistant might say. That is a model of a model: a confident metric backed by a guess. Teams then build content strategies on the estimate; the new content alters the clickstream, which feeds the next estimate, while the actual assistant may be giving completely different answers. A year of work can go into moving a prediction about a machine that nobody ever asked.
Even the tools that do query the assistants often blend the memory and live reads into a single composite score, which hides the diagnosis. And most platforms stop at the dashboard. In a July survey of 602 marketing and PR professionals by Scrunch and Scribewise, 46% said their organization may have moved too fast into AI visibility tactics without a clear long-term strategy. I think that is what it feels like to buy the dashboard before anyone has read it.
None of this estimating is necessary. An assistant will answer an automated query in the same full sentences it gives your customers. It costs more to query the models directly, which is why so many dashboards estimate instead, but it is the only way to build a remediation plan you can act on.
What to demand from any measurement
Hold any vendor, including my own company, to six requirements before you invest in AI measurement.
- Real questions. Test scores of the queries your customers actually type when they evaluate your category, not a handful of branded prompts.
- Real assistants. Run the queries on the exact free and paid assistants your customers use, and demand a precise list of which models were tested. Reject estimates and proxies.
- Separated reads. Report the memory read and the live read separately. Averaging them hides the fix.
- Receipts. Deliver the complete, unedited answer with the report. If an assistant recommended a competitor or made a claim about you, you need its exact reasoning and the source it cited, or you cannot defend the number when someone senior pushes back.
- A margin of error. Run the same query against the same model twice to see how much the answer moves. One run gives you a number; two independent runs give you a range, and next quarter’s change has to be larger than that range before anyone can call it real movement rather than noise. This is a different “twice” from item 3: there, each question is asked two ways, with search off and with search on; here, the same measurement is repeated to learn its own variation.
- Human accountability. A named reviewer who reads the data, says which gaps matter, and says what to fix first. A login to a dashboard is not that.
Jeff Dean, Google’s chief scientist, made a version of that last point in August. Describing how to make AI systems reliable, he suggested designs with “another model or another agent that’s evaluating which ones of those seem promising.” The system that produces an answer should never be the system that grades it. That holds for the assistants, and for anyone measuring them.
How Calafai approaches it
We built the Calafai Brand GEO Report to these specifications, so read what follows the way you would read any vendor grading its own homework.
The report takes scores of real customer questions and runs them against the assistants your buyers use: ChatGPT, Claude, and Gemini, each in its free and paid version, plus Grok, and Perplexity and Mistral in the single version available of each. A person on our team designs the question list from your market, your competitors, and the decision you are trying to make. Every question runs twice, once from memory and once with live search, and the two results are reported separately. We keep every raw answer and list every model tested.
What comes back is built to be acted on. For every question that matters, the report shows which assistants name you, which name a competitor or a retailer instead, their reasoning, and their sources, and it gives the fix, with live-read gaps first because those can be closed this quarter. Each fix names the page, the listing, or the outside source that has to change, so it can go straight to the internal owner. A senior brand strategist signs off on every report, and the executive summary usually arrives within two business days. From your side we need a short intake form and one person to hand the fixes to.
The report tells you what the assistants tell your customers, not what your customers think, so keep your human brand tracker running next to it. The margin of error in item five belongs to our scored measurement, Calafai Brand Standing, which settles what your brand should stand for in this channel and is the subject of the companion piece. Most brands need both, in that order: first decide what AI should say about you, then patch the places where it still says something else.
Two things an operations leader will want to know. We never recommend prompt injection (hidden text written to manipulate an assistant), synthetic content, or bot activity; those tactics work briefly and then get the whole domain penalized. And your intake data never trains an AI model, not ours and not those of the providers we query. Every question we run is a public question about a public brand.
None of this is assembled by hand. We run these engagements on the strategy software we built for our own work, which is how a measurement at this depth arrives in days. The software handles the volume, and people supply the judgment.
Who owns it inside your company
You cannot join the conversation before the click. No brand can. It happens one customer at a time, in a chat window you will never see. What you can do is decide what the assistant finds when it goes looking, and that takes four decisions inside your company.
- Create the row. Set up a standing quarterly measurement of what the assistants say to your top customer questions, with the memory and live reads kept separate. Until it appears on a slide someone presents, it does not exist.
- Assign one owner. Today this work fractures across brand, SEO, PR, and product. Pick a single accountable name.
- Fix the live-read gaps first. The sources the assistants cite are a finite list, and the teams managing them are already on your payroll.
- Audit your vendors. Before renewing any AEO subscription, make the vendor show that its numbers come from real answers, not clickstream estimates.
The conversation is happening either way, at a scale that already shows in Adobe’s traffic data and in your flat-traffic quarter. Find out exactly what is being said, and change it.
You cannot be in the room, but you can design the room.
If you would like to open a conversation about this, or learn what the assistants say about your brand today, write to [email protected].
Sources
- Immad Akhund, Founders in Arms, “Building Brokerage 2.0: Direct Indexing and Tax Alpha with Mo Al Adham,” August 14, 2026. The “direct money to the right place” quote.
- Abbas Haleem, Adobe: AI-referred traffic to retail sites doubles in a year, Digital Commerce 360, June 17, 2026, reporting Adobe Analytics data: AI-referred traffic to U.S. retail sites “grew 138% year over year in May 2026” and “converted at a rate 54% better than traffic from non-AI sources.”
- Bain & Company, Consumer reliance on AI search results signals new era of marketing, press release, February 19, 2025. “60% of searches now terminate without the users clicking through to another website.”
- Ara Kharazian, The a16z Show, “Why AI Isn’t Killing SaaS Yet,” May 25, 2026, originally aired on Monetary Matters. The “AEO” quote.
- Dominic-Madori Davis, AEO startup Profound hits unicorn valuation, raises $180M Series D 7 months after last round, TechCrunch, September 15, 2026. The round, the valuation, the previous round, and the company-reported revenue figure.
- Profound, Answer Engine Insights, product page, read September 17, 2026. The description of what the platform measures (“visibility score and share of voice,” “key themes, topics, and recurring narratives,” citations), how it runs (“We run every tracked prompt daily,” “We capture directly from the browser”), and that it “monitors and analyzes the responses generated by a RAG-based search approach.”
- Scrunch and Scribewise, 2026 AI search survey: Moving fast, flying blind, published July 16, 2026. 602 U.S. marketing and PR professionals surveyed May 19 to June 2, 2026. The 46% figure.
- Jeff Dean, Y Combinator Startup Podcast, “Jeff Dean: The 1% Rule for Building in AI,” August 1, 2026. The “another model or another agent that’s evaluating” quote.
Podcast quotations are taken from transcripts of the recordings and lightly trimmed with ellipses; nothing inside a quotation has been altered.