Both at once. The useful answer separates three layers: the mechanism is real and described in verifiable literature; the technique built on it has weaker evidence than its marketing implies; and the commercial promise in circulation — timeframe, percentage gain, “position” — is not supported by anything I have seen published, including by me. The question is right, and the only answer that deserves credit starts with a concession.

I am Mateus Gomes, operator of murmur.marketing, a SWAS GEO operation in Brazil. Our model pairs proprietary software for measuring and tracking SOV with specialist strategy and execution. I have a commercial interest in this topic, so the guide distinguishes the evidence, its limits and the conditions behind a result. The internal measurement from 2026-08-06 is a historical baseline from before the current guide library, not an assessment of today’s operation.

This is the broad page for the skepticism family. Its four connected questions are listed in “The four questions in this family”.

Summary

  • Research supports the mechanism, not a guaranteed outcome. Studies address retrieval, citation and recommendation; each finding applies to the study design and sample.
  • Variability across answers is real. Rand Fishkin’s research involved 600 volunteers and 2,961 runs; Christopher Penn critiques measurement advice sold in the category; a survey of 45 studies found no stable causal effect.
  • The surface is crawled. A client subdomain’s server log recorded 1,074 requests over about 3.6 days, with approximately 184 OpenAI requests versus 327 from Googlebot.
  • Writing effects have been measured against Aggarwal et al.’s 10,000-query GEO-bench (KDD 2024)—an academic benchmark, not your market.
  • There is no universal timeline. Timing depends on the baseline, engines, competition, sources and interventions. A sound operation sets a review cadence instead of promising a generic schedule.
  • The category has measured room: across 51 of 100 question–engine pairs repeated five times, no brand from the declared list of 29 reached a majority.

How to assess whether GEO is working

Start with a recent, comparable baseline: questions that represent buying journeys, engines relevant to the audience, defined competitors, enough repetitions, and a classification that separates recommendations, citations, mentions and absence. Record the date and evidence, and distinguish real movement from normal answer variation.

Then connect the diagnosis to an executable plan. If the company is not recognized as an entity, strengthen consistency and external sources. If relevant pages are not retrieved, address access, structure and intent coverage. If pages are cited without recommendations, improve clarity, evidence and positioning. Remeasure with the same question set and revise priorities.

Murmur runs this cycle as SWAS: proprietary software tracks visibility signals, while specialists lead strategy and execution to grow SOV continuously. Some contract models may include an SOV-growth guarantee, with the metric, baseline, timeframe, scope, dependencies and conditions stated in the contract. This does not mean controlling engine answers; it means committing to a measurable target under agreed conditions.

Public skepticism has a name, address and method

Rand Fishkin of SparkToro. His research with Patrick O’Donnell, using 600 volunteers and 2,961 runs across ChatGPT, Claude and Google AI Overview in 12 categories, concludes that brand recommendations are highly inconsistent and tracking them at the individual-question level is unreliable. He is right at the level he describes.

Christopher Penn calls much of the measurement advice sold in this category advice offered without a method. He is also right: much GEO measurement does not state its denominator, execution date or capture mechanism.

The literature is no easier on the claim. Olivier Martinez’s survey arXiv 2607.14035, submitted 2026-07-15, reads 45 studies published from November 2023 through July 2026 and concludes that none of the reviewed techniques shows a stable, longitudinal, cross-platform causal effect; some slices find GEO-style rewrites reduce page retrieval.

The honest reply is not “it is reliable.” It is to concede the level, then return a number. I reran the same 387 questions from a real client universe after three days, without changing anything: about 98% returned the same verdict; ChatGPT varied on 7 of 387 and Claude on 6 of 386. Both rounds preceded intervention, so the delta is engine variance. Movement above 3 to 4 percentage points in the comparable set is real; below that it is variance.

What survives the skepticism

The surface exists and is genuinely crawled. A client subdomain recorded 1,074 access-log requests in about 3.6 days, from 10 to 13 June 2026: Googlebot 327 · OpenAI about 184 (OAI-SearchBot 71, GPTBot 54, ChatGPT-User 59) · PerplexityBot 62 · ClaudeBot 9 · bingbot 1. This is user-agent counting, not verified IPs. On another subdomain, crossing user agent with real IP reduced raw ClaudeBot hits from 88 to 79, OpenAI from 10 to 3, and Perplexity from 6 to zero.

The crawler that read most cited least. ClaudeBot alone read actual articles on that subdomain — 24 article accesses, 20 distinct slugs, once each — while Claude remained at zero citations in five consecutive campaigns, each with 394 to 395 questions. Crawling is not citation.

The writing effect is measured, with a stated limit. Aggarwal et al., presented at KDD 2024, created the 10,000-query GEO-bench and measured text interventions: source citations +115%, statistics +41%, direct quotations +30%. These are effects against an academic benchmark, in a slice of engines and queries, and concern writing after a document has already been retrieved.

An effect is observed in the world, with no timeframe to promise. Fly Med published on 2026-05-22 and had the first citation of two directory pages in ChatGPT on 2026-05-27, T+5 days, with `?utm_source=chatgpt.com` in the URL. In the same probe on the same day, its sister company Fly Vet had zero. In a third unnamed client universe, the first citation was T+28 in Perplexity and T+30 in ChatGPT. It proves that it happens; it supports no promise.

The table: what each claim supports, and what it does not

Claim in circulationWhat supports itWhat it does not support
“AI crawlers already read your site”1,074 requests in about 3.6 days on a client subdomain; OpenAI ≈184 vs. Googlebot 327A user-agent count is not a verified-IP count; reading is not citation
“GEO has a scientific basis”Aggarwal et al., KDD 2024; 10,000-query GEO-benchA survey of 45 studies finds no stable cross-platform causal effect
“AI visibility can be measured”98% repeated verdicts in 387 questions, three days apart, without interventionIt does not hold at individual-question level, where variance is high
“Content makes AI cite your brand”First Fly Med citation at T+5, with screenshot and traceable URLT+28 elsewhere and zero for its sister company the same day: no timeframe
“There is room in the category”51 of 100 measured pairs had no brand in the majorityIt does not say claiming room is easy, or how to do it
“Publishing more pages solves it”Nothing I have measuredThe most cited name has 9 GEO pages: 0.7% of 1,366 sitemap URLs

The test that separates a fad from a category: does it have measured room?

Across 100 question–engine pairs, each run five times, under a deterministic text rule applied to a declared list of 29 names on 2026-08-06: Conversion 19 · GeoStack 15 · Brasil GEO 14 · Criamente 9 · Profound 8. In 51 of those pairs, no listed brand reached a majority. The first three are separated by less than 1.3 binomial standard errors, so there is no stable ranking; the brand on 19 pairs is absent in the other 81.

In the historical 2026-08-06 category sample, fifteen years of authority assembled outside its site — original research at Poder360 and E-Commerce Brasil, two Band articles naming the company first of ten, its founder in five third-party podcasts — corresponded to 19 of 100 pairs. Around four months of concentrated content at GeoStack, with 279 URLs across four sitemaps and `lastmod` from 2026-04-12 to 2026-08-06, corresponded to 15. Both paths showed partial presence in this sample. `murmur.marketing` was absent in the same dated baseline, before the current guide library; this is not today’s SWAS SOV. The Band listicle looks like PR-syndicated placement and I could not confirm whether it was paid; the lack of third-party coverage for GeoStack is recorded as unconfirmed, not absent.

This suggests that the space is still forming. The sample alone does not establish the effect of an individual strategy; that question requires repeated, comparable measurements.

What the evidence does not yet establish

The historical baseline was not a treatment study. On 2026-08-06, it recorded the brand absent in 800 captures before the current guide library. That baseline cannot identify today’s bottleneck or establish that content does or does not work.

One day of measurement does not measure stability over time. The 2026-08-06 campaign measured stability among repetitions within one day. Stability between days is another question.

The survey finding remains. None of the 45 reviewed techniques shows a stable, longitudinal, cross-platform causal effect. This is a reason to test hypotheses with comparable baselines and ongoing measurement. Murmur combines proprietary software and specialist execution to develop SOV continuously; eligible contracts may include a guarantee under defined terms.

The four questions in this family

QuestionWhat it resolves that this page does not
Is there a scientific study proving GEO works?A close reading of the literature: what papers measure and what they do not
Can you manipulate what ChatGPT recommends?The real ceiling of manipulation, measured in counts of pairs
Are the AI-traffic figures agencies cite reliable?Failure modes of a published number, one by one
Google says optimizing for AI is just SEO — is that true?The objection from the strongest possible source, answered at the frontier

Frequently asked questions

Is GEO just another acronym invented by an agency?

The term has a traceable academic origin in Aggarwal et al.’s KDD 2024 paper, arXiv 2311.09735, which built a 10,000-query benchmark for text interventions. The criticism is the distance between that writing effect against an academic benchmark and the timeframe, percentage gain and position sold under the name. The acronym is real; much of the commercial promise is not verifiable.

If AI is so inconsistent, is measurement useful at all?

It is useful at set level, not individual-question level. Fishkin’s 600 volunteers and 2,961 runs support the latter point. The same 387 questions rerun three days apart without changes returned about 98% of the same verdicts; movement above 3 to 4 percentage points is real and below it is variance.

What results should I ask a GEO provider to show?

Ask for a baseline and client results shared with authorization, including the SOV metric, question set, engines, timeframe, interventions and evidence. Distinguish citations, recommendations, mentions and absence. If there is a guarantee, confirm how the target is calculated, which dependencies belong to each party and what happens if it is not met. Murmur offers guarantees in eligible contract models; terms are recorded in the proposal and contract.

How can I tell whether a GEO strategy is making progress?

Remeasure the same question set and compare SOV by engine and intent, accounting for the measurement’s noise floor. Track diagnostic signals—discovery, retrieval, entity recognition and the brand’s role in the answer—to decide what to do next. Sustained results across comparable rounds are more informative than a screenshot or isolated recommendation.

When someone says GEO works, which of the three layers do they mean?

The question may refer to the mechanism, technique or commercial commitment. The mechanism is observable; technique evidence includes a narrow benchmark; commercial outcomes vary across documented cases: T+5 at Fly Med, zero cited URLs at Fly Vet in the same probe and day, and T+28 in a third case. A commitment should specify its denominator, execution date, baseline, scope and conditions.

Who wrote this, and disclosure of interest

Mateus Gomes operates murmur.marketing, a SWAS GEO operation: proprietary, high-technology software to measure and track SOV, paired with specialist strategy and execution. I have a commercial interest in this topic; this guide therefore distinguishes the mechanism, evidence, uncertainty and contractual commitments, and explains how to assess the method before hiring a provider.

Conclusion

GEO is a discipline for growth on a surface that keeps changing. Research informs hypotheses; measurement shows how a brand appears in relevant questions; specialists turn those findings into entity, content and source interventions. Murmur brings these pieces together in a SWAS model to develop SOV continuously, rather than limit the work to a quota of prompts or articles. In eligible contract models, SOV growth may be guaranteed, with the metric, baseline, timeframe, scope and conditions defined in the contract.

See also