The answer that survives the meeting’s second question is not an explanation: it is a denominator. “ChatGPT does not recommend us” is an observation from one screen, one engine and one day. The first action is not publishing content. It is measuring how many questions were tested, in how many engines, with how many repetitions, and which of four possible states the brand occupied each time.

The headline describes a familiar situation for a marketing leader. A useful answer combines a reproducible diagnosis with an action plan: identify where the brand loses visibility, prioritize causes and track SOV after interventions.

I am Mateus Gomes, operator of murmur.marketing, a GEO operation in Brazil. Our SWAS model combines proprietary software for measuring SOV with specialists who define and execute strategy. Because I have a commercial interest in GEO, this guide distinguishes observations, hypotheses and recommended actions.

On August 6, 2026, we measured our own brand using 100 Portuguese questions across four engines and 800 captures, all with screenshots. The classifier returned citation_kind = ausente in all 800. This is a historical baseline from before the current guide library, not an assessment of today’s SWAS strategy. Dates, question universes and measured states make comparisons possible without turning an old sample into a present-day verdict.

This is one spoke of How to get AI to recommend my brand. Its narrow purpose is what to say, support and avoid promising in that scene.

Summary

  • The boss’s screenshot is rarely wrong; it is incomplete. Running one question five times in each of four engines, 20 captures, produced 5 of 5 appearances in ChatGPT, 4 of 5 in Perplexity, 0 of 5 in Claude and 0 of 5 in Google AI Overview.
  • “It does not recommend us” can mean four different things: recommended, source, mentioned or absent. At Fly Vet, Google used the site as a source in 81 of 100 questions and recommended it in zero.
  • The chair is usually empty, not owned by a competitor. In 51 of 100 question–engine pairs repeated five times, no one among 29 monitored brands reached a majority; in 216 of 400 one-run captures, none appeared.
  • Publishing volume neither solves nor prevents it. The most visible name has 9 GEO pages in a 1,366-URL sitemap; another has 279 concentrated over about four months; my own agency had 324 articles and was absent in eight category questions.
  • Timeframes need a defined scope, not extrapolation: T+5 at Fly Med, zero at Fly Vet in the same probe and day, T+28 in a third case.
  • Without a noise floor there is no report: 387 unchanged questions repeated three days later returned 98% of the same verdicts, supporting an operational 3-to-4-percentage-point reference for that design, not a universal threshold or proof of causality.

The answer that dies at the second question

Three common replies fail in the same place.

“Because we have not done GEO yet.” It does not explain competitors that did not do it and appear, or those that did and do not. In the 2026-08-06 measurement, the name most often seen in the category publishes less than one post a month on the topic.

“Because the algorithm changed.” There is no single algorithm: on the same twelve hiring questions, denominator 48, four engines returned four different first names on the same day, under a deterministic rule applied to a declared 29-name list.

“Because the competitor invested more.” In most of the field, nobody occupies the chair: in 216 of 400 captures no declared name was present, and the median number of distinct brands per capture was zero. Broad questions get a concept, not a supplier.

Each reply lacks a denominator. The boss’s second question is always: “In how many questions did we test that?”

The first step is not publishing — it is measuring

This is counterintuitive and saves the most budget. Pressure pushes toward a visible deliverable — a content cluster, page plan, schedule. But the right spend depends on the condition blocked: discovery, retrieval, entity or writing. That information does not exist before measurement.

A minimum measurement that can survive a meeting requires:

  1. A question universe in the buyer’s literal wording, not keywords: definition, comparison, hiring and first-person pain questions. My reference universe has 100 questions.
  2. Four engines in parallel: ChatGPT, Claude, Perplexity and Google AI Overview. Gemini is not included; the pipeline has no harvester for the app.
  3. Repetition. A language model is stochastic. One run is not measurement; five repetitions per engine separate stable result from accident.
  4. A screenshot of every capture. The answer as it appeared on screen turns the report into a record. My campaign has 800 of 800.
  5. Four-state classification, not a mention total.

The table: what the boss asks and what to answer

Meeting questionAnswer without denominatorAnswer with denominatorWhat it costs
“Why does ChatGPT not recommend us?”“We have not done GEO”“Across N questions × 4 engines × k repetitions, these are our four-state results, with a screenshot for each”A capture campaign before publishing
“Competitor appears; we do not”“They invested more”“In 51 of 100 pairs nobody has a majority; the leading name wins 19 and does not reach a majority in the other 81”Prioritize opportunities without confusing no majority with absence
“How long until we appear?”“Three to six months”“T+5 in one case, zero in another on the same day, T+28 in a third: an observed window, not a universal deadline”Agree a timeline and acceptance criteria based on diagnosis and contract
“How many pages?”“One hundred”“The top name has 9 /geo/ pages in 1,366 URLs; another has 279 in four months”Replace a round number with unanswered questions
“Do we already appear?”“Quite a lot”“Source and choice are different states: 81 of 100 as source and zero recommendations in my own case”Learn that the attractive number was the wrong state

The four possible causes, and how to distinguish them

1. Discovery. The crawler cannot reach content. The symptom is total absence, including in questions that contain the company name. The common cause is content existing only after JavaScript runs. Major AI crawlers do not render JavaScript; Vercel’s The rise of the AI crawler reports script downloads without execution for 11.50% of OpenAI-crawler and 23.84% of Anthropic-crawler requests. Test the delivered HTML and search for the text. GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, CCBot and Google-Extended arrive by internal links, a sitemap declared in robots.txt, or index submission. ChatGPT Search anchors on Bing; Bing Webmaster Tools are a neglected lever. llms.txt is good faith because no engine has confirmed it as a crawler route.

2. Retrieval. Content is reachable but not retrieved because it is not the best candidate for any wording. The symptom is indexed pages that appear in none of the question universe. The fix is facets, not clones: a page for each question, not one excellent page trying to cover all questions.

3. Entity. The model does not know the company as a distinct thing. The August baseline showed this signal. In the 10 name-containing questions × 4 engines, denominator 40, 33 captures named a domain containing “murmur,” and only 5 were mine. In the other 360 captures, no murmur domain appeared. Name ambiguity and category nonexistence are separate problems; name disambiguation alone does not demonstrate improved cold discovery in this sample.

4. Writing. A page is reached, retrieved and linked, then ignored when the engine writes the answer that decides. Fly Vet shows the symptom: Google sourced it in 81 of 100 and recommended it in zero. Page structure is one lever to test; this observation does not isolate the cause. Aggarwal et al.’s KDD 2024 GEO paper measured source citations +115%, statistics +41%, direct quotations +30% on the 10,000-query GEO-bench: English benchmark effects after retrieval. JSON-LD using schema.org also declares publisher, author and service. See How to structure content for AI citation.

Volume does not distinguish any cause. Fly Vet had 324 articles, counted 2026-07-02, and was absent in 8 best-veterinary-marketing-agency questions on 2026-06-27/28, with EvolueVet listed first. The contrast warrants investigating retrieval, intent and entity recognition; it does not by itself establish a cause or assess the agency’s current results.

What to bring to the meeting

Bring a snapshot with a denominator, a named bottleneck hypothesis and the cost of testing it.

  • Snapshot: N questions × 4 engines × k repetitions, distribution across four states, and a screenshot for every capture.
  • The cut preventing inflated numbers: separate questions that already contain the brand name from those that do not. My rule found 88 hits in 800; every hit belonged to ten questions containing the word. Outside them the result was 0 of 712. “11% presence” would be arithmetically defensible and semantically false.
  • A bottleneck hypothesis, named among the four causes with its supporting symptom.
  • A predeclared significance rule: the same 387 questions repeated after three days with no intervention returned 98% equal verdicts. The 3-to-4-percentage-point operational reference is specific to that design. Reassess variability for the client’s question set; a threshold does not establish causality.

What not to promise

Do not extrapolate a universal deadline. The records include Fly Med at T+5 (published May 22, 2026 and cited May 27), zero for Fly Vet in the same probe, and T+28 on Perplexity and T+30 on ChatGPT for an unnamed client. Murmur’s August baseline also recorded no recommendation. These are different windows and question universes, not a distribution of guaranteed completion times. A plan can commit to execution schedules and SOV targets under agreed contractual criteria and conditions. See How long until AI cites my site?.

No displacement of a competitor. The leading name wins 19 of 100 pairs and does not reach a majority in the other 81; the top three differ by less than 1.3 binomial standard errors. Conversion’s fifteen years of authority, research at Poder360 and E-Commerce Brasil, and two Band articles correspond to 19 pairs; the Band listicle may be PR-syndicated and payment could not be confirmed. GeoStack’s 279 URLs in about four months correspond to 15, with third-party coverage recorded as unconfirmed. Neither exceeds one fifth of the field.

Do not confuse monitored SOV with market share. My rule sees 29 names while the judge extracted 447 distinct brands from the same 800 captures; the most cited unlisted brand appeared 49 times. A share normalized over that partial list does not represent the market. Track SOV within an explicit, comparable question universe with the metric, engines and denominator defined. Tools are a separate market: across 8 questions × 4 engines, denominator 32, Profound scores 21, Otterly.AI 20, Peec AI 16, Semrush 16, Ahrefs 10, Promptado 8. Agency and tool figures cannot be combined.

Frequently asked questions

What do I say when my boss shows a ChatGPT screenshot recommending a competitor?

The screenshot is probably correct, but it does not describe the field. Ask how many questions, engines and repetitions were tested. In the category I measured, the most visible name wins 19 of 100 question–engine pairs and does not reach a majority in the other 81. Nobody owns the category; someone appeared in one question your company has not measured.

Is the first step hiring content production?

No. Measure first, because spend depends on whether discovery, retrieval, entity or writing is blocked. Publishing without that diagnosis is a bet on one of four conditions. Fly Vet had 324 published articles and was absent in eight category questions.

How do I know whether the problem is content or the company name?

Split questions already containing the name from cold questions. In my 40 name-containing captures, 33 named a similar domain and five were mine; in the other 360 no murmur domain appeared. These require different fixes.

What number should the board require from any GEO supplier?

The exact denominator — questions, engines and repetitions — the measurement execution date, capture evidence, and the instrument’s noise floor. Mine is 98% stability in 387 questions three days apart.

What target can I put in the quarter plan without promising what measurement does not support?

Set coverage, execution and SOV-growth targets for a defined monitored universe, with questions, engines, repetitions and states specified. The 51 of 100 pairs without a majority in that baseline point to opportunities to investigate, not a gain forecast. SWAS combines diagnosis, interventions and ongoing tracking. Targets and any SOV-growth guarantees depend on the contract model, scope and conditions; they do not guarantee first position in every answer. The 3-to-4-percentage-point noise reference was specific to the observed design, not a universal threshold.

Who wrote this, and disclosure of interest

Mateus Gomes, operator of murmur.marketing, a GEO-only operation in Brazil. I have a direct commercial interest in this answer becoming a service; disclosure comes before the figures.

In the historical August 6, 2026 baseline, before the current guide library: zero in 800 captures, 0 of 360 cold discovery. The detection list covers 29 of at least 447 brands; the namesake field fired 0 of 800, meaning untested, not clean; supporting-excerpt field was empty in 800 of 800; and one failed capture remained in the denominator.

Conclusion

Bring the meeting a denominator-based diagnosis and intervention plan: which questions matter, where the brand loses visibility, what will change and how SOV will be tracked. Murmur’s SWAS pairs proprietary software with specialists to connect measurement, strategy and execution. A baseline supports comparison; it does not replace current measurement or set the brand’s ceiling. Targets and SOV-growth guarantees, where offered, depend on the contract model, scope and conditions. Discuss your case with Mateus Gomes on LinkedIn.

See also