Yes. To assess a GEO case, look for a starting baseline, interventions, methodology, timeframe and observed change in SOV—not just a screenshot of one favorable answer. Results vary by engine and question; a strong case shows the sample and its limits, including mixed outcomes, so readers can understand what was achieved and under which conditions.

I am Mateus Gomes, operator of murmur.marketing, a SWAS operation combining proprietary measurement software with specialist service to develop SOV. We publish evidence within the scope of documented authorization. This guide offers a method for evaluating any supplier’s case and distinguishing measurable impact from isolated examples.

This page answers existence. The selection question belongs to How to tell whether a GEO agency has real results.

Summary

  • A documented case needs a baseline, execution date, method and evidence that lets readers check the change.
  • Separate internal experiments from client outcomes. Fly Vet and Fly Med are businesses operated by the author; they illustrate variation, not independent client testimonials.
  • First-person pain questions gave Fly Vet 2 of 80 in ChatGPT and 0 of 80 in Claude.
  • Time varies by context. Observed cases can illustrate a range but do not establish a universal timeline or replace an agreed contractual target.
  • The sample is bounded. The 2026 campaign measured a declared set of questions and brands; it does not represent the whole market or establish current performance.

What must exist for a case to be documented

Denominator. “The brand began appearing” is meaningless without the total attempts: forty in fifty and forty in four thousand are different results.

Execution date, not publication date. Generative search changes between weeks. A case published now with three-month-old captures may describe a system that no longer exists.

Capture evidence. The answer as it appeared on screen allows checking without trusting the measurer and protects against silent instrument failure. Does a screenshot of an AI response prove a result? explains the boundary.

Limits and mixed results. A reliable case states what improved, what did not and which external factors affected the period. That helps buyers judge applicability without extrapolating one example to every brand.

What I can document by name: Fly Vet and Fly Med

They are companies I operate, so I can share their authorized figures with dates and denominators. That is the value and limit: these are internal cases, not independent client testimony.

Fly Med’s /geo/ directory was published 2026-05-22. The 2026-05-27 probe found two pages cited by ChatGPT, first and third in one question, with ?utm_source=chatgpt.com in the URL. In the same probe, same day and method, Fly Vet had zero cited /geo/ URLs. These are contrasting results from one dated probe, not a universal recipe or timeframe forecast.

The full probe has the cut almost nobody publishes: 32 mentions in 59 successful captures, of 65 attempted, equals 49% brand mention. Only two were /geo/ URLs. Mention and content working are different things.

In June 2026, 100 questions × one repetition under citation_kind: ChatGPT recommended Fly Vet in 30 of 100, Claude in 41 of 100, while Google AI Overview did not recommend it in that sample and used the site as a source in 81 of 100. A single combined index would erase the difference.

First-person pain questions (“my clinic does not appear on Google”) scored 2 of 80 in ChatGPT and 0 of 80 in Claude. EvolueVet appears in 28 of 100 ChatGPT answers; Fly Vet is recommended in 10 of those. In 8 best-agency questions Fly Vet is absent and EvolueVet appears first, despite Fly Vet’s published content cluster. These dated findings describe those question sets; they do not establish causality or today’s operational bottleneck.

There are seven first-place organic screenshots, gathered 2026-07-01, but their method is heuristic text position, not LLM judged. In the same collection, 3 of 11 direct comparisons show entity confusion.

The evidence table I can open

CaseMeasurementDenominator and dateResultWhat it does not prove
Fly Med — first citationChatGPT-cited /geo/ URLs2026-05-27 probe; publishing T0 2026-05-222 pages, positions 1 and 3, with UTM parameterThat timing repeats; its sister had zero cited URLs the same day
Fly Vet — same probeSame2026-05-27Zero cited /geo/ URLsIt is the counterfactual above
Fly Vet — mention × contentBrand mentions in successful captures32 of 59; 65 attempted; 2026-05-2749% mention; only 2 were /geo/ URLsThat mention came from content
Fly Vet — June campaigncitation_kind by engine100 questions × 1; 2026-06ChatGPT 30, Claude 41; Google sources 81 and recommends noneThat engines are comparable
Fly Vet — pain questionsFirst-person pain appearances80 records per engine, through 2026-06-30ChatGPT 2 of 80; Claude 0 of 80That category coverage covers pain across engines or contexts
Fly Vet × EvolueVetWho appears in category questions100 ChatGPT answers, 2026-06EvolueVet 28; Fly Vet 10 of those; Fly Vet absent in 8 category questionsThat published content is enough
murmur.marketing — 2026-08-06citation_kind on own brand800 = 100 questions × 4 engines; historical baseline before the current guide libraryAbsent in 800 of 800; 0 of 360 cold discoveryNot a current visibility result or an assessment of today’s SWAS strategy

What I cannot document, and why

The absence of a case is not absence of data; it is absence of permission. Some client-result material has no public-case consent artifact. I will not name a company, pair sector with size, or use a phrase that identifies it. Publishing performance without recorded consent is harm with a conflict of interest built in.

Anonymized material remains useful, with a privacy warning: the following records are from different materials and accounts; do not read them as one company.

  • The same 387 questions rerun after three days with no intervention repeated about 98% of verdicts: ChatGPT varied 7 of 387, Claude 6 of 386. Movement above 3 to 4 percentage points is real; below it is variance.
  • Removing an inflated claim across a client site changed nothing beyond that noise floor.
  • An unnamed client subdomain was first cited at T+28 in Perplexity and T+30 in ChatGPT.
  • ClaudeBot alone read actual articles in another log, while Claude had zero citations in five consecutive campaigns, each 394 to 395 questions. Crawler counts are user-agent, not verified IP.

Why the category has so few documented cases

In the historical 2026-08-06 campaign, denominator 400 — 100 questions × 4 engines — 216 named none of the 29 declared brands. The median number of brands per capture was zero. In broad-demand questions, denominator 120, the top name appeared 13 times. These dated category findings describe that sample, not the current SWAS strategy or present visibility.

Raw curl on 2026-08-06 saw Conversion’s sitemap with 1,366 URLs, only 9 under /geo/, and less than one GEO post per month since 2025-07-12. GeoStack had 279 URLs across four sitemaps, lastmod from 2026-04-12 to 2026-08-06, roughly four months, almost entirely on the topic. These are observed structures, not explanations.

Across 100 question–engine pairs in that historical campaign, each repeated five times under a deterministic rule applied to 29 names: Conversion 19 and absent 81; GeoStack 15; Brasil GEO 14; Criamente 9; Profound 8. The first three differ by less than 1.3 binomial standard errors. In 51 pairs no listed brand reached a majority. murmur.marketing appeared zero times in that 2026-08-06 sample, before the current guide library; this is not a current SOV measurement.

External context is evidence too

Olivier Martinez’s arXiv 2607.14035, submitted 2026-07-15, reviews 45 studies from November 2023 to July 2026 and finds no stable, longitudinal, cross-platform causal effect. SparkToro’s study, by Rand Fishkin and Patrick O’Donnell, used 600 volunteers and 2,961 runs. Aggarwal et al. at KDD 2024 (arXiv 2311.09735) built GEO-bench with 10,000 questions across 10 engines.

The google engine here is AI Overview. The Gemini app is not measured; no harvester exists for it.

Frequently asked questions

Does GEO have a documented success case in Brazil?

Documented examples include Fly Med’s two directory pages cited five days after publication and Fly Vet’s zero cited URLs in the same probe and day. These are dated internal agency experiments, not independent client testimony or clinic results. To assess applicability, check the baseline, period, interventions, questions, engines, denominator and capture evidence.

Does a case showing AI traffic prove GEO?

It proves a click, an optional consequence. A generated answer can solve a question without a visit, and a dashboard cannot say which brand the answer named. Traffic-only material measures a consequence, not the result.

Why do so few agencies publish GEO cases with numbers?

Recorded consent is necessary to publish identifiable client results. In the historical 2026-08-06 category campaign, 216 of 400 captures named none of the 29 listed brands and the median brand count was zero; this historical demand snapshot alone does not explain each provider’s publication choices.

Should a GEO case show varied results?

Yes. Favorable and unfavorable results with denominators and dates help estimate variance and show under which conditions an effect appeared. Here, the same directory led to a citation in five days at one company and zero cited URLs at its sister in the same probe and day; first-person pain questions returned 2 of 80 in one engine and 0 of 80 in another. These contrasts bound what can be concluded.

Is an anonymized case without a company name still evidence?

An anonymized case can be checked when it preserves its denominator, execution date, method and capture evidence. A company name enables direct confirmation, while the underlying arithmetic can still be checked without it. Records from separate materials should not be combined into a profile that risks re-identification.

Who wrote this, and disclosure of interest

Mateus Gomes, operator of murmur.marketing, a SWAS combining proprietary software to measure and track SOV with specialist strategy and execution. Fly Vet and Fly Med are agencies I operate; their data are internal examples, not independent client testimony.

For context, the 2026-08-06 baseline, before the current guide library, recorded the brand absent in 800 captures and 0 of 360 cold discovery. The campaign also recorded 2 of 80 and 0 of 80 for first-person pain questions, an intervention hypothesis with no movement beyond the noise floor, and a detection lens covering 29 names out of at least 447. These dated results disclose limits; they are not an assessment of today’s SOV.

Conclusion

Yes. Documented evidence exists, with dates, denominators and varied outcomes. Fly Med had two directory pages cited five days after publication; Fly Vet had zero cited URLs in the same probe and day. Other measured outcomes vary by engine and question family, while anonymized records include both movement and results within the noise floor. The murmur.marketing baseline of 2026-08-06 predates the current guide library and is not a current SOV result. Ask providers for baseline, scope, interventions, evidence and the conditions attached to any guarantee.

See also