Measuring AI visibility
EN

Measuring AI visibility

Last verified: July 1, 2026
13 min read
Case study
PageSpeed 100/100

#What we measured, and what we found

For a quarter we pointed AI-visibility monitoring at our own site and wrote down what it said, including the parts that were not flattering. The short version: our AI-citation rate is low where it matters most, not one of the three instruments we used read the product a customer actually opens, and one of our figures came out of a worksheet nobody filled in. The most useful output of the exercise was learning to distrust a clean-looking figure that came from the wrong instrument. The most humbling was finding out that for half the quarter we were the wrong instrument. This is the first of a quarterly series, so the numbers below are a baseline, not a victory lap.

Most writing about AI visibility is advice. This is measurement. We are an agency that argues for serving clean server-rendered HTML to AI, so it was only fair to check whether the agency itself gets cited. The answer, across three snapshots in April, May and June 2026, was more interesting than a single number.

#The instrument problem comes first

Before any finding, the caveat that reframes all of them. There are two common ways to measure how often an AI cites you, and they do not agree.

The cheap way is an API proxy. You send your prompts through a model’s API, read the text that comes back, and count brand mentions and links. It is repeatable and almost free, which is why it is everywhere. Its weakness is that the API path is not the product a customer uses. Published comparisons put the source overlap between API responses and the consumer web interface in the low single digits, so a proxy can tell you a model knows your name in the abstract while telling you nothing reliable about what a real user sees.

The honest way is to monitor the real consumer outputs, the same answers a person gets in the ChatGPT or Perplexity interface, including which sources the product actually surfaces. It costs more and is harder to automate. It is also the only number that corresponds to a lost or won visit.

We used three variants of the first and not once the second. The first draft of this report claimed otherwise, and that is the most reusable lesson in it: for an entire quarter we were measuring something other than what we said we were measuring, and we only caught it when we sat down to write out where each number had come from.

#April: the cheap proxy baseline

Our first snapshot, on 6 April, was an API-proxy run. The headline figures, across 26 queries against ChatGPT:

MetricResult
Brand-mention rate7.7 percent (2 of 26)
URL-citation rate0 percent
Strongest categoryPlugins, 14.3 percent mention
Transactional category12.5 percent mention
Informational and local0 percent mention

A mention rate under eight percent and a citation rate of zero reads like a disaster. The more useful detail is who got cited instead of us. When the model reached for a source on Polish WordPress work, it named directories and job boards: pracuj.pl appeared three times, alongside clutch.co, olx.pl, home.pl, nazwa.pl and a developer meetup listing. Those are not competitors who out-wrote us. They are aggregators the model trusts as generic answers to a commercial question. That pattern, the assistant defaulting to a directory rather than a specialist, turned out to be the real story, and on-page polish does not fix it.

The report itself carried the warning we want you to carry too: API proxy, directional trends only, roughly four percent source overlap with the web interface. We did not yet appreciate how much that warning mattered.

#May: the empty worksheet that produced a number

The 11 May run was not a run. The panel for that week is a worksheet in the repo: twenty queries split over ChatGPT, Perplexity, Bing Copilot and Claude, a blank result column beside each one, and an Owner: _______ line at the top that nobody ever signed. Nobody filled it in. The table our script produced from it, however, looked exactly like a measurement:

EngineRows in sheetFilled inReported result
ChatGPT60undetermined
Perplexity60undetermined
Bing Copilot40undetermined
Claude40undetermined

Every row came back undetermined, but not because the instrument checked anything and withheld judgement. scripts/compute-ai-citation-metrics.mjs reads a cell, counts it as a citation if it matches yes, y, true, 1 or x, counts it as no citation if it matches no, n, false or 0, and writes unknown in every other case. An empty cell matches nothing, so it gets unknown. Twenty undetermined results are twenty rows nobody looked at. There is no grounding detection in that pipeline, and there never was.

The last line of the report is worse. The citation rate is computed as cited-yes divided by all rows, so those twenty empty rows went straight into the denominator, and the generated file announces: citation rate 0 percent. An empty worksheet went into the script and came out the other side as a neatly formatted report with a zero in the headline and a four-engine breakdown underneath.

That is what this failure mode looks like from the inside. Not an instrument honestly reporting that it cannot see. An instrument converting missing data into a result and presenting it with the same confidence as a measurement. We accuse AI-visibility dashboards of this a few paragraphs from now, so it is only fair to say it plainly: our own pipeline did exactly that, nobody noticed, and May’s zero sat in the repo looking like a finding the whole time. A number does not have to be invented to mislead. It only has to be well formatted.

#June: a second instrument, a second blind spot

On 11 June we changed instruments. We described this to ourselves at the time as moving from the API to real model outputs. It was not. We connected Geoboard, six models including ChatGPT, Claude, Gemini, Perplexity, Grok and DeepSeek, against a set of buyer-intent prompts. That is a different vendor, not a different kind of measurement: our client for it, scripts/geoboard-pull.mjs, is a dozen lines wrapping curl around a REST API. There is no consumer interface anywhere in it.

Geoboard’s own methodology note says the thing we were not saying: the run queries models without live web search, which makes it a measure of training-data recall. A real ChatGPT and a real Gemini search live and would find the running site. So the June run was not measuring whether AI cites us. It was measuring whether the models remember us from training, which for a niche B2B agency is close to zero by construction. We had traded one blind spot for another and briefly mistook it for progress.

The picture sharpened anyway, because two blind spots in different places show more together than one does alone. For one narrow query, a Polish studio serving foreign WordPress clients, we ranked first in five of the six models tested. With live search disabled that means something specific, and stronger than we had assumed: the association sits in the models themselves, not in search results stapled on at query time. It is a specific identity claim with little competition, exactly the kind of query a specialist should own.

For transactional WooCommerce queries, the queries closest to revenue, we were close to invisible. The models answered confidently and did not reach for us. ChatGPT was consistently the weakest channel for the brand, without a single hit across the tested set. Perplexity came out strongest.

And here we nearly made the exact mistake this report is about. The first draft of this paragraph explained Perplexity’s lead by noting that Perplexity leans hard on live web search. The explanation is tidy, it fits everything known about the product, and it cannot survive here: no model in that run had live search. Whatever gave Perplexity the lead, it was not live search. We do not know what it was. Guessing would have turned a method caveat we already had in writing into a sentence that sounds like a conclusion, which is the same move as May’s zero, just better dressed.

The competitors who did surface were mostly general SEO and SEM agencies marketing themselves as doing “AI SEO”, not WooCommerce specialists. The gap, in other words, is authority and association, not page quality.

#What the three snapshots add up to

Read together, the quarter says three things plainly.

First, the method is part of the number. The three runs are spread across roughly nine and a half weeks, from 6 April to 11 June, and the site did not hold still in between: we shipped new six-locale decision guides in that window aimed squarely at the gap the measurement had exposed. So we cannot say the thing we would like to say, which is that the method moved the number rather than reality. We have no control for that, because reality was moving too. We can say something weaker and sounder: these three methods diverge so sharply that no figure among them is interpretable without the method that produced it. If a tool gives you an AI-citation rate without telling you whether it read the API or the product, and whether the model had live search on, distrust it.

First-party measurement is the point of this whole exercise, and it is the same discipline we apply to client work: a number you cannot reproduce is not evidence. We learned the same lesson the hard way with a synthetic-brand experiment, where a model confidently described a company that did not exist. Measuring a real brand has the opposite failure mode, confidently reporting zero when the truth is unknown, and our May script did that literally. Both come back to the same obligation: check the instrument before you trust the reading, including, and especially, when you wrote the instrument yourself.

Second, identity queries are winnable and transactional queries are not, at least not on-page. We hold the narrow positioning query because it is specific and lightly contested. We lose the commercial queries because the models lean on directories and broad agencies, and no amount of cleaner HTML changes who a model already associates with “WooCommerce developer”. That is an off-page authority problem.

Third, the channel matters, even when you cannot yet say why. The same brand came out strongest in Perplexity and did not appear in ChatGPT once, in the same run, on the same query set, under the same settings. We have no explanation for that. What we do have is the per-model breakdown that shows the difference, and that is enough to know where to look next. A single blended “AI visibility score” hides exactly the information you need, and throws in a false sense that you understand something.

#What we changed

We did not rewrite pages in response to a proxy number, because that would be optimising for an instrument rather than a customer. Instead, the measurement changed where we spend effort.

  • We fixed the query set and the cadence: a stable list of identity, informational and transactional queries, recorded monthly, with the engine, the date and the signature of whoever filled the row in. The Owner line stopped being decoration.
  • We label every number with its instrument and its settings, including whether the model had live search, and we do not put figures from two different instruments side by side.
  • May stays in the series marked as no data, not as zero. The scoring script still counts empty rows in the citation-rate denominator, and until we fix that, no zero out of that pipeline goes into a report without someone checking how many rows were actually filled in.
  • We moved the transactional-visibility work off-page, because the gap there is association and authority, not on-page content, and that is documented in our off-page authority plan rather than in another rewrite.
  • We kept serving everything in server-rendered HTML, which is the precondition for being citable at all and the subject of our note on why Western assistants read raw HTML.

#How to run this yourself

You do not need our budget to start. The minimum honest setup is a fixed list of ten to twenty queries that real customers would ask, run once a month, with three columns recorded every time: the engine, the date, and whether your brand was named or your URL linked. Add a fourth column for which other domains were cited, because that tells you who you are actually competing against in the answer, which is rarely who you think. Add a fifth for who filled the row in. It sounds like bureaucracy right up until you see a report generated from a sheet nobody touched.

If you use an automated tool, ask it two questions before you trust a single chart: are you reading the API or the product, and did the model have live search on? If it cannot answer, treat the output as directional only, the way all three of our runs were. And if the tool is your own, check one more thing: what it does with an empty row. Never let it turn missing data into a confident zero.

#The honest takeaway

A quarter of measuring our own AI citations produced one uncomfortable number, our transactional citation rate is low, and one genuinely valuable habit, never quote an AI-visibility figure without the method that produced it. May did not show us at all, it only looked like it had. June surfaced a defensible identity position April could not see, and did it while measuring something other than what we thought at the time. All three readings became useful only once we wrote down how each was taken. Before that, one of them was just a zero in a file. This is report one. We will publish the next snapshot at the end of the quarter, on the same query set, so the series can be compared rather than admired. If you want to be cited by AI, start by measuring it honestly, and built that into the workflow we describe for GEO and LLMO.

Next step

Turn the article into an actual implementation

This block strengthens internal linking and gives readers the most relevant next move instead of leaving them at a dead end.

Want this implemented on your site?

If visibility in Google and AI systems matters, I can build the content architecture, FAQ, schema, and internal linking needed for SEO, GEO, and AEO.

Related cluster

Explore other WordPress services and knowledge base

Strengthen your business with professional technical support in key areas of the WordPress ecosystem.

What is an AI-citation rate and why measure it?#
It is how often an AI assistant names your brand, or links your URL, when it answers a relevant question. It matters because AI answers increasingly sit between a searcher and your site. If the assistant cites a directory instead of you, you lose the visit before classic SEO even applies. You cannot improve what you do not measure, so the first step is a baseline.
Why did the three runs produce such different numbers?#
Because they were three different instruments, and none of them read the product a customer uses. The April run went through a model API, and its own note puts source overlap with the web interface at a few percent. The May panel was a worksheet nobody filled in, so its zero was never a measurement. The June run queried six models through a vendor REST API with live search switched off, so it measured training memory. Those are three different questions, not three readings of the same quantity.
What did the June run actually show?#
A split, though a narrower one than we first assumed. For the narrow identity query about a Polish studio serving foreign WordPress clients we ranked first in five of six models. With live search disabled that means the association sits in the models themselves. For transactional WooCommerce queries, where the money is, we were close to invisible. ChatGPT was the weakest channel for us and Perplexity the strongest, but we do not know why, because no model in that run had live search.
How often should I measure AI citations?#
Monthly is enough for a small site, because model behaviour and your own content both move slowly relative to the measurement noise. Pick a fixed query set, record the engine, the date and the settings with every figure, and never put numbers from two different instruments side by side. Record who filled the row in, too, or you cannot tell absence of a citation from absence of a measurement. We publish a quarterly snapshot so the series stays honest and comparable.
Is a high AI-citation rate worth chasing for every query?#
No. Identity and informational queries are easier to win and worth holding, but transactional queries are where competitors and directories fight hardest, and where on-page work alone rarely moves the needle. For those, off-page authority matters more. Measure first, then spend where the gap is real, not where the number is easy.

Need an FAQ tailored to your industry and market? We can build one aligned with your business goals.

Let’s discuss

Related Articles

Why Perplexity cites your brand and ChatGPT does not

Our own Geoboard baseline showed Perplexity as the strongest engine and ChatGPT with zero presence across eight tracked prompts in the same run. Here is the mechanism behind that split, and what it means for procurement, evaluators, and agencies reporting AI visibility to clients.