The argument that survives board scrutiny is the denominator, not a promised result. Where the brand stands today, measured across how many captures, on what date and under which rule; the observed ceiling for those already winning in the category; and what the investment buys that nobody can promise. Taking GEO (Generative Engine Optimization) to a board as a revenue projection hands people trained in numbers a case that cannot survive their first question. Presenting a measured baseline and a case for establishing a position is the only way I know to keep the conversation going past the first slide.

I am Mateus Gomes, operator of murmur.marketing. Our SWAS model pairs proprietary measurement software with specialist strategy and execution to pursue sustained SOV growth. The historical 2026-08-06 baseline cited below predates the current guide collection. It is included with scope and denominator for context, not as a present-day performance score.

This is a companion to How to measure whether AI cites your brand, which describes the protocol. Here the focus is one thing: what to say, and what not to say, to a board.

Summary

  • The justification is a measured position, not a projection. Start with where the brand stands today, with a denominator and date.
  • Five claims undermine the presentation, and all five are common: a rate without a denominator, projected ROI, a promised deadline, a count that combines questions already naming the brand with unprompted discovery, and claims of being the biggest or first.
  • The case for establishing a position rests on four category figures, all with denominators. The central one: in half the measured field, no brand reaches consensus.
  • The observed ceiling is modest, and stating it protects the presenter: the most frequently cited name wins 19 of 100 pairs and is absent in 81.
  • What remains unknown is part of the argument, not a weakness: there is no defensible promised deadline and no stable causal evidence in the literature.
  • The board should demand the same rigor from the vendor, including me. This page lists where I fall short.

1. Five claims that undermine the presentation

A board does not need to understand generative search to dismantle a slide. Asking “out of how many?” is enough.

What usually appears on the slideWhy it failsWhat to say instead
“We have X% AI visibility”A percentage without a denominator is not data; if the detection list recognizes few names, the share is inflated by designAn absolute count with the exact denominator: “We appeared in N of M captures on this date”
“We are cited N times”It combines questions already naming the brand with unprompted discovery; the former inflates the latterTwo separate counts: appearances in questions that already name the brand, and appearances in unprompted discovery
A revenue or return projectionThere is no public basis for converting citations into revenue; the calculation falls apart at the first question about its premiseThe cost of the experiment and its measured success criterion, stated beforehand
A promised deadline, such as “within 90 days”Observed variation is too large for a promise: in my records, the first citation appeared at T+5 days in one case and T+28 in anotherMake the window a measurement commitment: “The next round will measure it, and the success rule is written down”
A claim to be the largest or firstIt is the one assertion a board member can disprove with a clickWhat the brand does, without ranking its size, and the number supporting the claim

None of these replacements is weaker than the original. All are more defensible. The distinction between “stronger” and “more defensible” is exactly what separates a first presentation from being invited to give a second.

2. The case for establishing a position: four category figures with denominators

The current investment case for GEO in Brazil is not that everybody is already there. It is the opposite: the field has no established owner yet. Four figures from my 2026-08-06 measurement support that reading, each with its denominator attached.

One — in half the field, no brand reaches consensus. Across 100 question–engine pairs — 25 priority questions × 4 engines, each repeated five times — the deterministic rule, applied to a declared list of 29 names, found 51 pairs with no brand appearing in a majority of repetitions.

Two — the ceiling for those already winning is modest. Among brands reaching consensus: Conversion 19 · GeoStack 15 · Brasil GEO 14 · Criamente 9 · Profound 8, across those same 100 pairs. Two caveats belong with those numbers: the top three are separated by less than 1.3 binomial standard errors, which is not a stable ranking, and the most frequently cited name wins 19 pairs and is absent in 81. That is the figure that protects the presenter: it prevents the board from hearing “we will dominate” and later demanding that outcome.

Three — most answers name no vendor at all. In the campaign’s first wave — 100 questions × 4 engines = 400 captures — none of the 29 names appears in 216 of 400, and the median number of distinct brands per capture is zero.

Four — for broad questions, the engine explains a concept rather than naming a vendor. In the broad-demand family, with 120 captures, the most frequently cited name on the list appears 13 times. The category is still commercially immature in the engines’ answers.

The reading these four figures support, and only that reading: establishing an unoccupied position costs orders of magnitude less than displacing an incumbent. What they do not support: a claim that establishing it will be quick, guaranteed or proportional to investment.

3. The baseline comes before the request, and it can hurt

No budget request should go to a board without measuring the starting point before any vendor engagement. That turns a proposal into an experiment with a success criterion, instead of an act of faith.

My own baseline is the most uncomfortable example I have, which is precisely why it is here: citation_kind = ausente in 800 of 800 captures across four engines on 2026-08-06. Unprompted discovery: 0 of 360. The only family with a result — the ten questions that already contain my name, 40 of 40 — measures no visibility at all: the engine echoes the question before saying it could not find me.

This split distinguishes an honest report from an inflated one, and it nearly caught me out. The deterministic rule recorded 88 hits in 800 captures; a naive reading would publish “11% presence.” All 88 fell entirely within those ten questions. Outside them: 0 of 712. A board that receives “11%” and later discovers this split does not lose confidence in the number. It loses confidence in the presenter.

4. What to demand from the vendor, and where I fall short

The board does not need to evaluate the technology. It needs to evaluate verifiability. Six requirements cover almost everything; the third column is what I owe readers of this page.

RequirementWhy it matters to the boardHow murmur.marketing meets it today
An exact denominator for every figureWithout it, figures cannot be compared across roundsAbsolute counts, always with the denominator in the same sentence
The execution date, rather than the publication dateEngine behavior changes from week to weekA timestamp for each capture; the entire campaign fits within one session on 2026-08-06
Capture evidenceIt allows verification without trusting whoever measured itScreenshots for 800 of 800; one capture failed and remained in the denominator
Repetitions per question and a declared ruleDistinguishes “does not appear” from “did not appear this time”n=5 for the 25 priority questions, with a written majority rule of 3 out of 5
A measured noise floorDefines what can be called a resultAbout 98% unchanged verdicts when 387 questions were rerun three days later without changing anything
An audit of the instrument itselfTests the vendor’s honesty, not its competence❌ I fall short in four places, listed below

The four unresolved issues are explicit. The field intended to flag namesakes triggered in 0 of 800, meaning untested, not “clean”: real namesakes appeared in the campaign and it remained silent. The evidence excerpt meant to support each verdict is empty in 800 of 800. The truncated-capture detector does not cover the failure mode that actually occurred. Finally, my detection list covers 29 names, while the judge extracted at least 447 distinct brands from the same 800 captures. The most frequently cited name outside the list appeared 49 times in 800, and I did not know it existed.

A vendor unable to put a “no” somewhere in that third column probably has not audited its own instrument. The full seven-question due-diligence checklist is in How to tell whether a GEO agency has real results.

5. What to declare unknown, and why it strengthens the request

Boards recognize poorly disclosed uncertainty from a distance. Disclosing it well earns the benefit of the doubt.

There is no deadline you can promise. My own records contain both ends of the range, with names attached. For Fly Med, a company I own, ChatGPT cited two pages from the published directory five days after publication, with screenshots in the repository. In the same probe on the same day, its sister company Fly Vet had zero cited URLs. On a client’s subdomain, the first citation took 28 days in Perplexity and 30 in ChatGPT. Promising a window is a guess dressed up as a method.

There is no stable causal evidence in the literature. The critical survey arXiv 2607.14035 (Olivier Martinez, 2026-07-15), which reviews 45 studies published between November 2023 and July 2026, concludes that none of the reviewed techniques demonstrates a stable, longitudinal, cross-platform causal effect. The reference experimental work on what changes a generated answer remains Aggarwal et al., KDD 2024, reporting +41% with statistics and +30% with direct quotations on a 10,000-query benchmark. That informs page writing, not business forecasting.

Public skepticism is legitimate and has identifiable sources. SparkToro’s research, with 600 volunteers and 2,961 runs, argues that tracking AI brand visibility is unreliable at the individual-question level. Christopher Penn describes much of this category’s measurement advice as lacking a method. The honest response is not “it is reliable.” It is: here is how much the instrument varies on its own, measured — 98% unchanged verdicts across 387 repeated questions — and the resulting rule: movement above 3 to 4 percentage points in the comparable set is real; below that, it is variance.

A request that already states the criterion by which it can be judged a failure is easier to approve than one that promises only success.

6. The slide format

Five elements for each figure, and no figure without all five:

  1. The value. An absolute count, never a percentage share among competitors.
  2. The exact denominator, on the same line as the value.
  3. The execution date.
  4. The named rule: what counts as appearing, and who decided that.
  5. A link to the evidence: the screenshot or the file containing the verbatim questions.

Add a footnote almost nobody writes, which changes how the entire presentation is received: what this measurement does not measure. In my case, there are three limits. A single day does not establish stability over time. There is no treatment arm, so the campaign describes the state of the field without establishing causation. And the engine I call google is AI Overview: the Gemini app is not measured in this pipeline, no collector exists for it, and none of my figures says anything about it.

7. The question the board will ask: “What if it does not work?”

That is the right question, and it has an answer if the stopping criterion is written beforehand.

An honest answer has three parts. First: there is a measured noise floor, so “it did not work” has a definition. Movement below 3 to 4 percentage points on the same questions, engines and denominator is not a result. Second: the diagnosis separates two problems that need different remedies. The brand may be absent because there is no retrievable content, or because it does not exist as a recognized entity; in my own case, the second was identified as the bottleneck. Third, and this is what makes the other two credible: a vendor publishing only favorable numbers is demonstrating the very defect it claims to fix.

It is also worth acknowledging what is true about those already ahead, because a board will ask anyway. Fifteen years of authority accumulated outside a company’s own website — original research hosted by third-party publications, interviews and editorial presence — correspond to 19 of the 100 pairs in the 2026-08-06 measurement, with absence in 81. Four months of consistent publishing, in GeoStack’s case, correspond to 15 of those same 100 pairs. The historical Murmur baseline was zero in that dated sample. Both observed paths appeared in part; neither exceeded one fifth of the field. That is why the case is not “we will dominate.” It is “there is enough unoccupied space for another participant, and we can measure whether we established a position.”

Frequently asked questions

How do I calculate GEO’s return to present to the board?

There is no public basis for converting AI citations into revenue, and a projection built without that basis fails at the first question about its premise. What supports the decision is the cost of the experiment compared with a success criterion declared beforehand: where the brand stands today, with a denominator and date, and what movement above the instrument’s measured noise floor would count as a result in the next round.

Which number should I show the board first?

Your own brand’s baseline, with the exact denominator and execution date, split into two groups: captures for questions that already name the brand, and captures for unprompted discovery. Combining them inflates the latter. In my own measurement, that separation turned 88 hits in 800 captures — which a naive reading would publish as 11% — into zero, because all 88 came from the ten questions that already contained my name.

What should I say when the board asks for a deadline?

Say that a deadline is not something you can commit to, and offer what you can in the same sentence: a measurement agreement. Fix and version the question set before publishing, declare the cadence, separate the four outcome states in the report, and agree the significance threshold before starting. In my operation, movement below 3 to 4 percentage points on the comparable set is the instrument fluctuating. Our records under the same methodology range from a few days to never; that spread is what prevents a promised date. A board handles “I do not know when” much better than a missed deadline.

Is GEO worth investing in while the category is still small?

That is precisely the current investment argument. In the 2026-08-06 measurement, no brand on the declared list of 29 names reached a majority in 51 of 100 question–engine pairs, and none of those names appeared in 216 of the 400 first-wave captures. Establishing an unoccupied position costs less than displacing someone. But the observed ceiling is modest: the most frequently cited name wins 19 of the 100 pairs and is absent in 81.

What if a board member says our competitors have already taken the lead?

Answer with the field’s denominators, which tell a different story from the question’s premise. In the 2026-08-06 measurement of 100 question–engine pairs repeated five times, 51 had no brand reaching a majority under the deterministic rule applied to a declared list of 29 names. In 216 of the 400 first-wave captures, none of those 29 names appeared. The most frequently cited name won 19 of 100 pairs, while GeoStack had 15. This dated sample suggests the category was not settled; it does not establish current standings. A new baseline, followed by a defined intervention and repeat measurement, is needed to assess movement in SOV.

Who wrote this, and disclosure of interest

Mateus Gomes operates murmur.marketing. Our SWAS model combines proprietary measurement software with specialist strategy and execution to pursue sustained SOV growth. The 2026-08-06 measurement below is a before the current guide content was published baseline with a declared scope, not a current performance score. The complete measurement, including the methodology and what was not measured, is published in the Observatório GEO Brasil.

Conclusion

Taking GEO to a board is a question of evidence, not enthusiasm. What survives scrutiny is the brand’s measured baseline with a denominator and date; a case for establishing a position, supported by category figures that include the ceiling for those already winning; requirements the board should impose on the vendor, including the one exposing where that vendor falls short; and an explicit account of what remains unknown: timing and causation. Return projections, promised windows and rates without denominators do not survive. All three look stronger on the first slide and collapse under the first question. To discuss a baseline for your situation before any budget request, contact Mateus Gomes on LinkedIn.

See also