Measuring whether AI cites your brand requires six decisions, five of them made before the first capture: which questions, how many repetitions, which engines, whether to capture through the browser or API, which classification to use, and — the decision almost nobody makes — what the instrument’s noise floor is. Without the last one, any fluctuation becomes a “result.”

I am Mateus Gomes, operator of murmur.marketing, a Brazilian GEO operation combining proprietary measurement software with specialist strategy and execution to grow SOV. The protocol below supports a cycle of baseline, diagnosis, intervention and follow-up measurement. Its historical example predates the current guide collection. If the distinction between GEO, AEO and SEO is still an open question for you, see GEO, AEO and SEO: what is the difference?.

Protocol summary

  1. A fixed, versioned question set. Written beforehand, with the reason for each question alongside it. Questions that change between rounds measure nothing comparable.
  2. Repetition. A single run is not measurement. Repeat, at a minimum, the questions that matter.
  3. Four engines by default. ChatGPT, Claude, Perplexity and Google AI Overview disagree with one another. Measuring one tells you about one.
  4. Browser captures with screenshots. This is a different interface from the API.
  5. Four classification states, not a boolean: recommended, source, mentioned and absent.
  6. A measured noise floor before calling any fluctuation a result.

Two rules of honest reporting matter just as much as the protocol: an exact denominator for every figure, and never combine questions that already name the brand with unprompted discovery.

1. The question set

The question set is the instrument. If it changes between measurements, there is no comparison: there is a sequence of snapshots of different things.

Three practical requirements:

  • Fixed and versioned. A dated file in a repository. The next comparison runs the same questions.
  • A written reason beside every line. Why was that question included? Without that record, months later nobody knows what the figure means. In my set, every line has a rationale field written before the campaign; afterward, each also has a result. One becomes the content brief, the other the priority.
  • Split by role. Buying questions (such as “what is the best company that does X”), method, tool, skepticism and brand questions measure different things and cannot be summed into one figure.

The right size depends on what you want to detect, and there is a point beyond which more questions stop helping. Too many dilute the signal and consume repetitions, the scarce resource. I prefer a hundred questions, with five runs for the priorities, to a thousand questions run once.

2. Repetition: why a single run is not measurement

This is where most AI visibility reports fall apart.

In my 2026-08-06 campaign, the question that started the project — “qual a melhor empresa pra fazer GEO no Brasil” (English translation: “what is the best company for GEO in Brazil”) — was run five times in each of the four engines, producing 20 captures. The company named in the screenshot that prompted me to start was Conversion. Here are the results of the 20 captures, using the deterministic rule against a declared list of 29 names. The measured question remained in Portuguese:

EngineConversion named, out of 5 repetitions
ChatGPT5/5
Perplexity4/5
Claude0/5
Google AI Overview0/5
Total9/20

The screenshot that started all this was not wrong. It was incomplete: in one engine the company always appears; in two others it never appears, across the same five repetitions on the same day. A screenshot of one question in one engine on one day is anecdotal evidence, including when it favors the person presenting it — especially then.

This behavior is not peculiar to my instrument. SparkToro’s research, conducted by Rand Fishkin and Patrick O’Donnell with 600 volunteers and 2,961 runs in ChatGPT, Claude and AI Overview across 12 categories, measured the same inconsistency at a much larger scale. It concludes that tracking brand visibility at the individual-question level is unreliable. I agree with the finding and disagree with the operational conclusion often drawn from it: the response is repetition and denominators, not abandoning measurement.

My rule for saying a brand “wins” a (question, engine) pair with five repetitions is a majority: appearing in at least 3 of 5. It is a choice, it is declared, and any reader can recalculate under a different rule.

3. Browser or API: the decision that changes the number

This is not an implementation detail. They are different interfaces, with different routing, grounding and search tools. What the API returns is not what a person sees when asking the product.

The strongest evidence is not mine: a vendor selling API-based measurement published a comparison of 1,000 prompts across OpenAI’s API and chatgpt.com, finding 96% divergence. Someone with a commercial interest in defending the API measured and published the opposite finding.

Browser measurement costs more: clean sessions, interface automation and screen capture. In return, it delivers what the API cannot: a screenshot. That turns a number into something a client can verify without trusting me.

For the 2026-08-06 campaign, screenshots exist for 800 of 800 captures, or 100%. The files exist on disk and were checked individually.

What “verifiable” means, side by side

We can make this concrete without passing judgment on anyone’s work. The Portuguese-language AI citation study closest to mine is the Citation Share Report — Bancos Digitais no Brasil (“Digital Banks in Brazil,” translated title), published by GeoStack and read by me on 2026-08-06. Comparing the two only by what each document states in its own text:

DimensionCitation Share Report (GeoStack)My 2026-08-06 campaign
Observations1,000 — 200 prompts × 5 platforms800 — 100 questions × 4 engines
Repetitionsn=1 per platform, stated in the textn=1 for 100 questions and n=5 for 25
Execution dateNot stated; publication date is 2026-04-292026-08-06, recorded per capture
Capture mechanismNot stated: neither browser nor API is specifiedBrowser with the interface, explicitly stated
Screenshot of the answerNone; aggregated tables onlyScreenshots for 100% of captures
Questions retained verbatimNoYes, in a versioned file, available on request

This is not a judgment about that study’s quality or the company that published it. It lists what each document declares and omits, which any reader can verify by opening both. A study can be correct without being verifiable; it simply cannot be checked by someone who did not run it. The difference is not who is right, but what each document lets the reader check independently. As someone who sells measurement, that is the standard I accept being held to.

4. Classification: four states, not a boolean

A brand can occupy four states within an answer, each requiring a different remedy. The labels below are English translations of the Portuguese measurement categories:

StateWhat happenedRemedy
RecommendedPresented as a choice in the bodyMaintain the position and expand coverage
SourceLinked, but absent from the bodyPage writing
MentionedNamed, without becoming a choicePositioning
AbsentOutside the candidate setDiscovery or entity recognition

A single “mentions” or “win rate” metric merges the first three and returns the same number for opposite problems. Always ask for a breakdown by state. The two most often confused, “source” and “recommended,” are examined individually in Cited as a source or recommended by AI.

Two independent rules. I classify the same captures in two ways: a deterministic text rule declared in a versioned file, and a model-based judge that reads the entire answer. The first is reproducible and blind to context: it knows that a word appeared, not what the sentence says about it. The second reads the sentence and makes different errors, including missing passing mentions. I never mix them in the same table. When they disagree on the order of names but agree on the structural finding, that is information. Combining them in one table would make it mere confusion.

Here is the concrete example from my campaign. Across the same 100 (question, engine) pairs, each repeated five times, the text rule found 51 pairs with no brand reaching a majority. The judge’s rule, applied to those same 100 pairs, found 55. They disagree on the order of names and agree on the finding that determines the interpretation: in roughly half the field, no brand reaches consensus. That is all the comparison supports. It does not establish which ranking is true; it establishes that the structural finding does not depend on which rule you choose. That is why the two rankings never appear together anywhere in this guide.

5. The noise floor: the step almost nobody takes

Measure the noise before selling the signal. Generative search varies between identical runs. Without knowing how much it varies on its own, any fluctuation can be read as an effect of an intervention.

What I did: I ran the same 387 questions from a real client’s set three days apart, changing absolutely nothing between rounds. About 98% of the questions returned the same verdict. In ChatGPT, 7 of 387 changed; in Claude, 6 of 386. Both rounds preceded any intervention, so the difference is pure engine variance.

This yields the significance rule I use, rather than an opinion: movement above 3 to 4 percentage points in the comparable set is real; below that, it is variance. A vendor celebrating a 2-point rise is celebrating noise — and cannot know otherwise without measuring the floor.

That number is also a direct response to identifiable public skepticism, which I prefer to cite rather than loosely paraphrase. Drawing on SparkToro’s research with 600 volunteers and 2,961 runs, Rand Fishkin argues that AI brand-visibility tracking is unreliable at the individual-question level. Christopher Penn goes further, describing much of the measurement advice sold in this category as lacking a method. The critical survey arXiv 2607.14035 (Olivier Martinez, 2026-07-15), reviewing 45 studies from November 2023 to July 2026, concludes that none of the reviewed techniques demonstrates a stable, longitudinal, cross-platform causal effect.

All three are right about the state of the market. The honest response is not “it is reliable.” It is: here is how much the instrument varies on its own, measured, and the rule that follows. Conceding the point at the individual-question level does not settle it at the level of the set: 387 repeated questions returned 98% unchanged verdicts. Without measuring your own noise floor, you cannot know which of those two levels you are discussing.

6. Two distinctions that separate honest figures from inflated ones

Distinction 1 — never combine questions naming the brand with unprompted discovery

These are two experiments. One measures whether the engine recognizes a name already in the prompt; the other measures whether it discovers the brand on its own. Combining them lets the first inflate the second.

In my case, the deterministic rule recorded 88 hits in 800 captures for my brand. A naive reading would publish “11% presence.” All 88 came from the ten questions that already contained the name. They were echoes of the prompt, in answers saying the opposite of recognition. Outside those ten questions: 0 of 712.

In the historical 2026-08-06 baseline, citation_kind = ausente (the original “absent” label) appeared in 800 of 800 captures across all four engines. This distinction was exactly what separated publishing “11%” from a correctly scoped discovery result. The campaign predates the current guide collection and is not a present-day visibility score. It matters just as much in client reports, where getting it wrong is more costly.

Distinction 2 — never publish a normalized share among competitors

A text rule sees only names declared in a list. Mine contained 29 names. The judge, which does not use a list, extracted 447 distinct brands from the same 800 captures: 418 outside my list. The most frequently cited unlisted brand appeared 49 times in 800, and I did not know it existed.

The mechanical consequence: any share calculated across those 29 names is inflated, because a brand’s share of a universe of 29 is larger than the same share in a universe of 447. That is why I publish absolute counts with a declared denominator, never percentage shares among rivals. When a report says “12% share of voice,” ask: twelve percent out of how many names?

Here is how I publish that information in practice. These are absolute counts of captures, out of 400 — 100 questions × 4 engines, one repetition each — in which a name appears in the answer body. They use the deterministic rule on the declared list of 29 names, dated 2026-08-06:

BrandCaptures, out of 400
Conversion50
GeoStack49
Brasil GEO39
Profound37
Peec AI32
Otterly.AI30
Semrush29
HubSpot25
Promptado22
Ahrefs21
murmur.marketing0

Two things about this table, and the second matters more than the first. First: it includes the Murmur pre-publication baseline, which was zero in this dated campaign; that is historical context, not a current result. Second: it must carry the context it hides on its own. None of the 29 names appears in 216 of the 400 captures, and the median number of distinct brands per capture is zero. This is not a field contested by a short list of giants; most answers contain none of them. That is why none of these figures becomes a percentage here.

What this protocol does not cover

Four limits, stated so readers can hold the method to them:

1. One day of measurement does not establish stability over time. The 2026-08-06 campaign measured stability between repetitions within one day. Stability between days is different, and only a subsequent round measures that.

2. There is no treatment arm. A snapshot describes what the engines answered on a particular day. It does not establish what makes a brand improve. Confusing those two things is the most common error when reading a visibility report, and the problem extends beyond my case. The critical survey arXiv 2607.14035 reviews 45 studies and concludes that none of the reviewed techniques demonstrates a stable, longitudinal, cross-platform causal effect. A serious measurement protocol does not solve that; it simply stops pretending to have solved it.

3. The instrument has a declared scope. The engine I call google is AI Overview, the generated block at the top of search results. The Gemini app is not measured: this pipeline has no collector for it, and none of my figures can be read as evidence about it.

4. An automatic detector cannot settle namesake cases. In my campaign, the field meant to flag “this is another company with the same name” remained at 0 of 800. That means untested, not “clean”: the campaign contained real namesakes, and the field remained silent throughout. For any brand with an ambiguous name, a positive verdict must be confirmed by reading the screenshot. The detector cannot make that decision.

One further limit concerns the purpose of the instrument: measuring is not fixing. This protocol identifies the brand’s state and stops there. What to do with that diagnosis is covered in How to appear in ChatGPT as a company, which describes the order of the four conditions. In my own case, the identified bottleneck was the third: entity recognition.

Frequently asked questions

Can I measure this with an off-the-shelf tool?

Yes. Check three things before buying: whether it captures through the browser or API, whether it repeats each question, and whether its report separates the four states. The browser/API distinction is not cosmetic: a vendor selling API-based measurement found 96% divergence between OpenAI’s API and chatgpt.com across 1,000 prompts, a finding I collected on 2026-08-02. A tool returning a single mentions count without a denominator or capture evidence cannot support an investment decision.

What do I measure if my brand does not appear anywhere?

Measure the field, not only your own brand. Even when a brand is absent from a sample, a report can show which questions have an established winner, which do not, which domains engines use as sources, and whether the pattern is uniform across engines or concentrated in one. In Murmur’s historical 2026-08-06 baseline, the dated result still helped identify where to investigate first. A measured result is most useful when it leads to a defined intervention and a follow-up SOV measurement.

What should I do when two rules disagree about the same capture?

The disagreement is a finding, not a defect to resolve by choosing one rule. I run a deterministic text rule, which is reproducible and blind to any brand outside its declared list, alongside model-based extraction, which sees unlisted names but is inconsistent with passing mentions. When they agree, the verdict is straightforward. When they disagree, that capture goes to human review. A collection of precisely those disagreements revealed that one count was inflated by echoes of the prompt. I never combine the two in one table: they answer different questions.

Do I really need screenshots?

A screenshot lets the client check the result without trusting whoever measured it. It is also the only defense against a silent capture failure: one of my 800 captures failed because the engine built an interactive artifact instead of answering, and I discovered it only because there was a file to inspect. I deliberately kept that capture in the denominator. Excluding individual captures invites cherry-picking what counts.

What counts as a real improvement?

Movement above the instrument’s own measured noise floor, across the same questions and engines with the same denominator. In my case, the floor came from running 387 questions twice, three days apart, without changing anything: about 98% of verdicts were unchanged. Movement above 3 to 4 percentage points is real; below that, it is variance. Improvement on questions already naming the brand is not evidence of improved unprompted discovery. They are different mechanisms, and attributing one to the other teaches the wrong lesson.

What does Search Console’s Generative AI report show?

Available worldwide since August 31, 2026, Search Console’s Generative AI performance report shows impressions, pages, countries, devices and trends over time for Google AI features in Search, including AI Overviews and AI Mode. Discover has a separate generative report. The tool does not show answer text or prompt-level data, so it does not replace answer capture or measure SOV comparably across engines. Google Search Central and Search Console Help.

Who wrote this, and disclosure of interest

Mateus Gomes operates murmur.marketing, a Brazilian GEO operation combining proprietary measurement software with specialist strategy and execution to grow SOV. The protocol is intended to connect a verifiable baseline to action and follow-up measurement.

The 2026-08-06 campaign is a historical example from before the current guide collection: citation_kind = absent in 800 captures across four engines. The correction separating brand-naming prompts from unprompted discovery illustrates why the question universe and denominator must be explicit. This result is not a current visibility score. The full measurement, including methodology and what was not measured, is published in the Observatório GEO Brasil.

Conclusion

Measuring whether AI cites your brand is a protocol, not a query: a fixed, versioned question set, repetitions, four engines, browser captures with screenshots, four classification states, and a measured noise floor before calling any fluctuation a result. Two rules of honest reporting support the rest: an exact denominator for every figure, and separation between questions already naming the brand and unprompted discovery, the distinction that separated “11%” from zero in my own case. There are no percentage shares among competitors, because those depend entirely on how many names the detection list recognizes. To discuss measurement for your situation, contact Mateus Gomes on LinkedIn.

Five questions in this topic group

This page describes the full protocol. Each question below addresses one specific decision within it, has its own page, and links back here. Published pages are linked; any unlinked item is still in the queue.

QuestionWhat it resolves beyond this page
Measuring GEO through the API or the browser: which to trustWhere to capture, using seven methodological criteria
How to build a question set to test my brand in ChatGPTWhich questions belong, their roles, and why exact wording matters
How many questions you need for reliable AI measurementThe size of the set and repetition count, with a measured noise floor
I’m a CMO taking GEO to the board — how do I justify it?The same measurement in the language of a budget decision
How often to measure AI visibilityA cadence derived from the noise floor

See also