Seven questions separate measurement from a pretty report, and all seven can be answered before signing a contract: the exact denominator for every number, execution date (not publication date), capture evidence, repetitions per question, the declared scale, a measured noise floor, and a versioned question universe. A provider that answers all seven can still be wrong; one that answers none cannot even be checked — and that distinction is the only one a buyer can assess from outside.
I write as Mateus Gomes, operator of murmur.marketing. Our SWAS model pairs proprietary measurement software with specialist strategy and execution to pursue sustained SOV growth. This guide offers due-diligence criteria for any provider, including Murmur. The historical 2026-08-06 baseline predates the current guide collection and is not a present-day score.
This is the broad page in the evidence family. Its six hanging questions are mapped in “The six questions in this family”.
Summary
- The right question is not “do you have a case study?” It is “what is the denominator?” A case study without a denominator is an anecdote with a logo.
- Seven criteria, all verifiable before signing: denominator · execution date · capture evidence · repetition · declared scale · noise floor · versioned universe.
- A screenshot of one question in one engine on one day is not evidence. In my campaign, the same question returned a brand in 5 of 5 ChatGPT repetitions and 0 of 5 Claude repetitions, on the same day.
- A useful proposal connects four parts: diagnosis, strategy, execution and ongoing SOV measurement. Evaluate the work and its terms, not just dashboard access or content volume.
- Fact against fact, without adjectives. The AI-citation study closest to mine in Portuguese declares 1,000 observations (200 prompts × 5 platforms) and “each prompt run once per platform”; it declares no execution date or capture mechanism and provides no screenshot. A study can be correct without being verifiable — those are different things.
- Nobody in the category has a comfortable position. In 51 of 100 (question, engine) pairs run five times, no brand in the declared list of 29 reaches a majority.
The seven questions, and why each one hurts
1. What is the exact denominator?
“The brand appeared 40 times” means nothing. How many attempts? In how many engines? With how many repetitions? A number without a denominator is an opinion dressed as data, and the asymmetry favors the seller: 40 in 50 is one result, 40 in 4,000 is another, and the sentence is the same.
The rule I apply and recommend demanding: put the denominator in the same sentence as the number. Not in a footnote, appendix, or methodology note at the end of a PDF.
2. What is the measurement execution date?
Not the publication date. Generative search changes behavior between weeks, and a report published in July with March data describes a system that no longer exists. The uncomfortable question — and the one that separates documents most — is: “On what day were these captures made?”
3. Is there capture evidence?
A screenshot makes verification possible without trusting the measurer. It is also the only defense against silent failure: in my campaign, one capture in 800 failed — Claude built an interactive artifact rather than answering — and I only knew because there was a file to inspect. I deliberately kept that capture in the denominator: excluding isolated captures invites cherry-picking what counts.
There is a technical consequence almost nobody discusses: browser and API are not the same surface. The strongest evidence is not mine — a provider that sells API mode published a comparison of 1,000 prompts between the OpenAI API and chatgpt.com and found 96% divergence between them. The party with a commercial interest in defending the API measured and published the opposite. If a report does not say which surface it captured, it does not say what it measured. The detail is covered by GEO measurement by API or browser: which should you trust?.
4. Was every question run more than once?
This is where most reports come apart. In my 2026-08-06 campaign, the question that started the project — [English translation of the Portuguese prompt] “what is the best company for GEO in Brazil?” — ran five times in each of four engines, 20 captures. Under a deterministic text rule over a declared list of 29 names, Conversion was named in 5 of 5 ChatGPT repetitions, 4 of 5 Perplexity repetitions, 0 of 5 Claude repetitions and 0 of 5 Google AI Overview repetitions.
The screenshot that made me start all this was not wrong. It was incomplete. A screenshot of one question, in one engine, on one day, is an anecdote — including when it favors the person showing it, especially then. Third parties have measured this inconsistency at greater scale: Rand Fishkin and Patrick O’Donnell’s SparkToro research, with 600 volunteers and 2,961 runs in ChatGPT, Claude and Google AI Overview across 12 categories.
5. What is the declared scale?
“Mentions” is a word, not a scale. The minimum usable classification separates four states — recommended, source, mentioned, absent — because they require different remedies and one metric merges the first three. A report returning one number for “visibility” returns the same number for opposite problems. The scale is open in Cited as a source or recommended by AI.
Also ask for the winning rule: where there are repetitions, what is the cutoff? Mine is a majority — the name appears in at least 3 of 5 repetitions. It is a declared choice, and every reader can recalculate it differently.
6. Did the provider measure its own noise floor?
Measure noise before selling signal. Without knowing how much an instrument moves on its own, any variation can be read as intervention effect. My floor came from running the same 387 questions from a real client universe three days apart, without changing anything: about 98% returned the same verdict — 7 of 387 fluctuated in ChatGPT and 6 of 386 in Claude; both runs preceded any intervention, so their delta is pure engine variance.
This produces a significance scale that is not opinion: movement above 3 to 4 percentage points in the comparable set is real; below that, it is variance. A provider celebrating a two-point rise is celebrating noise — and cannot know otherwise because it did not measure the floor.
7. Is the question universe fixed, versioned and public?
If questions change between measurements, there is no comparison — only a sequence of photographs of different things. Ask for the file. Ask for the reason for every line, written before the campaign. Ask for questions verbatim: a universe that only exists described in prose (“purchase and comparison questions”) is not auditable.
Due diligence: seven criteria for comparing GEO operations
Use these criteria to assess any provider, including murmur. The goal is not to find an operation without uncertainty; it is to understand what it measures, what it executes and how it responds to results.
| criterion | what to ask for | why it matters |
|---|---|---|
| Business objective | target SOV, category, engines, questions and timeframe | Connects GEO work to a goal that can guide priorities. |
| Baseline | versioned question set, date, denominator and recommendation/citation scale | Makes comparable measurement possible across rounds. |
| Diagnosis | gaps in entity clarity, content, sources and competitive visibility | Explains where the brand loses ground and where work can help. |
| Execution | deliverables, owners, cadence and client dependencies | Distinguishes applied strategy from dashboard access or a task list. |
| Monitoring | repeated measurement, evidence and priority reviews | Shows movement and informs adjustments. |
| Commercial model | limits on content, prompts, engines, usage and support | SaaS can be efficient, but plan limits should fit the goal and be explicit. |
| Guarantee | metric, baseline, term, scope, obligations and contractual remedy | A guarantee is useful when its terms can be checked. |
SaaS, services or SWAS: compare the delivery model
A SaaS tool can provide autonomy and scale; some plans also cap prompts, projects, users or content. Check each provider’s current plan and contract: limits vary and change over time. Do not assume every SaaS product has the same restrictions, or that raising a quota alone will increase SOV.
murmur operates as SWAS (software with a service): proprietary measurement technology combined with specialists who set priorities and implement GEO interventions. The offer is not just a dashboard or a bundle of articles. It connects diagnosis, strategy, execution and monitoring to pursue sustained SOV growth. Scope is designed around the agreed strategy rather than treating a generic content or prompt quota as the outcome.
Under certain contract models, murmur may offer a guarantee of SOV growth. This is not an unconditional promise: eligibility and terms depend on a baseline, question universe, engines, timeframe, measurement method and each party’s responsibilities, all set out in the contract. Ask for those details and the remedy if the target is missed in writing.
How murmur turns measurement into action
Proprietary software tracks responses across an agreed question universe; the team interprets where the brand appears, in what role (recommended, source, mentioned or absent), and against which alternatives. It then prioritizes interventions such as clarifying the entity, strengthening pages that answer real questions, improving structure and evidence, and building credible sources about the brand. New measurements inform the next round. This is an optimization process, not control over third-party answers.
The historical campaign dated 2026-08-06 covered 100 questions across four engines and 800 captures. It describes that baseline, not murmur’s current performance or a forecast for clients. Commercial decisions should use a recent baseline and the method defined in scope.
What public category documents declare — and do not declare
This is fact against fact, and the comparison is what each document says about itself, not anyone’s quality.
The AI-citation study closest to mine in Portuguese is Citation Share Report — Bancos Digitais no Brasil, published by GeoStack and read by me on 2026-08-06. Its own text declares 1,000 observations, 200 prompts × 5 platforms; [English translation of its Portuguese wording] “each prompt run once per platform”; and classification through API entity extraction with manual review. It does not declare an execution date — publication was 2026-04-29 — whether capture was browser or API, or provide answer screenshots.
| dimension | Citation Share Report (GeoStack) | murmur.marketing campaign, 2026-08-06 |
|---|---|---|
| observations | 1,000 — 200 prompts × 5 platforms | 800 — 100 questions × 4 engines |
| repetition | n=1 per platform, declared in text | n=1 in 100 and n=5 in 25 |
| execution date | not declared (published 2026-04-29) | 2026-08-06, stored per capture |
| capture mechanism | not declared | declared browser interface |
| answer screenshot | none | screenshot in 100% of captures |
| questions verbatim | not retained | retained in a versioned file, supplied on request |
This is not a judgment of that study’s quality or of the company that published it. It is a list of what each document declares and does not declare, checkable by any reader opening both. A study can be correct without being verifiable — it simply cannot be checked by someone who did not execute it. Because I sell measurement, this is the line on which I accept scrutiny.
The evidence base should be part of any provider review. The arXiv 2607.14035 survey by Olivier Martinez, submitted 2026-07-15, reads 45 studies published between November 2023 and July 2026 and concludes that no reviewed technique demonstrates a stable, longitudinal, cross-platform causal effect. Christopher Penn argues that much measurement advice in this category is offered without a method. These caveats set limits on what providers should claim; they do not replace an evaluation of the measurement, strategy, execution and contractual terms offered by each provider.
The 2026-08-06 historical sample is a useful example of why buyers should resist treating a single ranking as a verdict. Under a deterministic text rule over the declared list of 29 names, in the same 100 (question, engine) pairs run five times, Conversion appeared in 19 and was absent in 81; GeoStack, 15; Brasil GEO, 14; Criamente, 9; Profound, 8; and in 51 of 100 pairs no listed brand reached a majority. The first three are separated by less than 1.3 binomial standard errors, so this is not a stable ranking. The Murmur baseline in this before the current guide content was published campaign is historical and should not be read as a current result. Fifteen years of authority beyond one’s own site and a concentrated content footprint are two distinct patterns to investigate, not causal findings from this sample.
The six questions in this family
This page establishes the criterion. Each question below is a closed part of it with its own page and points back here. Those already live are linked; the two without links depend on a client data slice that has not yet been authorized, which is also declared.
| question | what it resolves that this page does not |
|---|---|
| Who in Brazil publishes original GEO measurement data | the map of who publishes figures and under what declared method |
| Does an AI-answer screenshot count as evidence of results? | what one screenshot proves, what it does not, and how to falsify it |
| A Brazilian study of which brands AI recommends | the criteria an AI-citation study needs to declare |
| Does GEO have a documented success case in Brazil? | the mapped absence in the category, answered with available evidence |
| Is there a real case of a company that began to be recommended by ChatGPT? | ⏳ depends on a client data slice not yet authorized |
| Which GEO agency shows client before-and-after results? | ⏳ depends on the same slice |
Frequently asked questions
What question should I ask a GEO agency first?
“What is the denominator?” Before the case study, portfolio or price. Any AI-visibility number — mentions, citations, recommendations — only means something against total attempts, engines used and execution date. A provider presenting “40 mentions” without saying 40 out of how many has usually not hidden anything deliberately; it simply did not measure in a way that permits the question.
How do I know whether the agency measured my category or only my name?
Ask for the question universe verbatim and count how many prompts already contain your brand. A question naming the company measures recognition; a cold-discovery question measures whether you exist in the category, and neither replaces the other. In my material, 88 matches for my name in 800 captures were ready to become an 11% headline, but all 88 fell in the 10 questions that already contained the word — outside them, 0 of 712. A document that does not separate those families cannot tell you which number you bought.
How can I compare two GEO agencies that present different numbers?
Do not compare numbers; compare what each document declares: exact denominator, execution date, capture mechanism, repetitions per question, classification scale and versioned question universe. If one declares all six and the other two, the comparison is not results but a checkable document against one requiring trust. A study can be correct without being verifiable; it cannot be checked by someone who did not run it.
What is a noise floor, and why does it matter when hiring?
It is how much a result moves by itself, with no intervention, when the same measurement is repeated. Without it, any variation can be sold as work effect. Mine came from running the same 387 questions from a real client universe three days apart without changing anything: about 98% of verdicts repeated. That yields the scale that movement above 3 to 4 percentage points is real and below it is variance. A provider that has never measured its floor cannot know whether its chart is work or noise.
How should I evaluate an SOV-growth guarantee?
Ask for the metric and baseline definition, the question universe and engines, the evaluation window, each party’s dependencies, and the remedy if the target is missed. Compare the same method at the beginning and end, and check how scope changes or engine availability are handled. murmur offers SOV-growth guarantees under certain contract models and conditions; the proposal and contract should spell them out.
Who wrote this, and the declaration of interest
Mateus Gomes operates murmur.marketing, a GEO operation in Brazil. As a provider, he has a commercial interest in this subject; this guide therefore makes its buyer criteria explicit and applies them to every proposal. murmur combines proprietary measurement software with specialized strategy and execution services to grow SOV in AI answers.
Conclusion
A GEO operation should connect goals, measurement and execution. When comparing providers, ask what SOV they will track, how they will establish the baseline, what actions they will take, how tool limits affect the plan and what terms govern any guarantee. Murmur combines proprietary high-technology software with specialized services to pursue sustained SOV growth; eligible contract models may guarantee growth under agreed conditions. The proposal should make scope, method and terms explicit so you can decide whether the model fits your brand.