Brazilian studies of which brands AI recommends do exist. What separates a usable number from a decorative one is not its result, but nine things the document does or does not disclose about itself: execution date, denominator, capture mechanism, repetitions per question, capture evidence, verbatim questions, classification rule, the size of its measurement lens, and conflicts of interest. This page lists those nine disclosures in the order in which they change how a result should be read, comparing what public documents in this category disclose and omit.
I will disclose my interest before giving any figures. I am Mateus Gomes, operator of murmur.marketing, a GEO operation in Brazil combining proprietary measurement software with specialist strategy and execution. I have a commercial stake in this category, so this guide applies the same methodological criteria to every study and separates what the data state from what they cannot support. The internal measurement from August 6, 2026 is a historical baseline from before the current guide library, not an assessment of today’s SWAS operation. It remains a useful example of why results need a date, denominator and context.
The parent page is how to tell whether a GEO agency gets results. That page assesses a supplier. This one explains how to read a study.
Summary
- The nine criteria are disclosures, not results. A study that meets all nine can still be wrong; one that supplies only two cannot be checked.
- The first criterion is the execution date, not the publication date. That distinction already separates the available documents.
- The classification rule changes the number: a mention, a citation and a recommendation are different events. A single “visibility” score merges them.
- The hidden measurement lens matters: a list of 29 names versus a judge that found 447 in the same 800 answers, including 418 outside the list.
- In a real family of 48 hiring answers, the engines disagree about the leader, and 8/48 answers contain none of the 29 names.
The nine criteria, in the order in which they change the interpretation
1. Execution date — not publication date
Engines change within weeks. A report published in July using captures from March may describe a system that no longer exists in that form. This is the cheapest information to request and one of the most frequently missing.
In GeoStack’s Citation Share Report — Bancos Digitais no Brasil, as read on August 6, 2026, the publication date is April 29, 2026; the execution date is not disclosed. In our campaign, all 800 captures carry individual timestamps from August 6, 2026, in a single session.
2. The exact denominator, in the same sentence as the figure
“The brand appeared 40 times” means something very different out of 50 answers than out of 4,000. A publisher benefits when the denominator disappears. It should not be buried in a footnote, an appendix or a methods section at the end.
Our denominator is 800 captures: 100 questions across four engines, asked once for the full set and five times in total for a subset of 25 questions. A number without its denominator is not yet a comparable measure.
3. Browser or API
This distinction is not academic. A vendor that sells API-based monitoring tested 1,000 prompts through the OpenAI API and chatgpt.com and reported 96% divergence — a result against its own commercial interest. If a study does not disclose its capture mechanism, the reader does not know what it measured.
The buyer uses an interface. An endpoint is not automatically a measurement of that interface.
4. Repetitions per question
Language models are stochastic. GeoStack’s report states “cada prompt rodado uma única vez por plataforma” — “each prompt run once per platform” (translation from the original Portuguese). That is n=1, explicitly disclosed, not a hidden flaw. The reader should simply not interpret it as a stability test.
In our campaign, the Portuguese prompt “qual a melhor empresa para fazer GEO no Brasil” — “what is the best company for GEO in Brazil” (translation of the measured prompt) — was asked five times per engine, for 20 captures. Within the deterministic list of 29 names, Conversion appeared in 5/5 on ChatGPT, 4/5 on Perplexity, 0/5 on Claude and 0/5 on Google AI Overview. A single run at a different moment could have produced four different reports.
5. Capture evidence
A screenshot lets another person verify the answer without relying on the publisher’s interpretation. It also reveals silent failures: in one of our 800 captures, Claude produced an interactive artifact instead of an answer. We noticed it by inspecting the file. That capture remained in the denominator; removing an inconvenient failure would have selected the sample after seeing the result.
There is a dedicated page on whether a screenshot of an AI answer counts as evidence.
6. Verbatim questions — and discovery versus echo
Saying in prose that a study used “purchase and comparison questions” is not enough to audit it. Ask for the file, with the exact text of every question on its own line.
A brand already named in a prompt tests recognition or echo, not discovery. Across the 800 captures, there were 88 text hits for “murmur.” All came from the 10 branded questions. Outside those questions, the result was 0/712 in the 2026-08-06 baseline, before the current guide library. It is a historical sample, not a measure of today’s SWAS strategy or SOV. Combining echo and discovery in one measure changes the meaning of the result.
7. A written classification rule
“Mentions” is a word, not a rule. At minimum, four states need to be distinguished: recommended, cited as a source, mentioned, and absent. They call for different actions. A single score merges the first three. See being cited as a source versus recommended by AI.
The rule for winning across repetitions also needs to be stated. We use a majority of 3/5. A reader who disagrees can recalculate under another declared rule.
8. The size of the measurement lens
This is the least visible source of inflation. A text-based measurement using a fixed brand list defines the universe it can see. Our list contained 29 names. The judge found 447 distinct names across the same 800 answers: 418 were outside our list. The most frequent outside name, Wyse, appeared in 49/800. We did not know that name when we assembled the questions.
Normalizing competitors’ shares within a universe of 29 instead of 447 makes the shares larger. “12% share of voice” — out of how many names? Publish absolute counts alongside their denominators.
9. Disclosure of interests
Studies in this category may be published by vendors of a solution, including ours. That does not invalidate them. It does require disclosure before the figures, together with the publisher’s own result and its date. Our 800/800 absence finding belongs to the 2026-08-06 baseline, before the current guide library; it is not an assessment of today’s SWAS operation or SOV.
When an author ranks well in an undisclosed study of its own category, the preceding eight criteria become more important, not less.
The table: what each document discloses and omits
This comparison reflects the documents as read on August 6, 2026. It is not a quality ranking. A correct study can be impossible to verify; a missing disclosure does not make a result false, but prevents the reader from checking it.
| Criterion | GeoStack — Citation Share Report | murmur.marketing — August 6, 2026 campaign | What the disclosure lets a reader check |
|---|---|---|---|
| Execution date | Not disclosed; publication date: April 29, 2026 | August 6, 2026, timestamped per capture | Whether the measured system still exists in that form |
| Exact denominator | 1,000 = 200 questions × 5 platforms, disclosed | 800 across 100 questions × 4 engines, with the disclosed repetitions | Magnitude, rather than direction alone |
| Capture mechanism | Not disclosed | Browser interface, disclosed | Which interface was actually measured |
| Repetitions per question | n=1, disclosed | n=1 for all 100; n=5 for a subset of 25 | Stability versus a single draw |
| Capture evidence | No screenshots published | 800 out of 800 | Independent manual reassessment |
| Verbatim questions | Not published | Published in a versioned file | A comparable future measurement |
| Classification rule | API entity extraction with manual review, disclosed | Four states and a 3/5 majority rule, disclosed | What the count means |
| Measurement lens | Not disclosed | 29 names out of at least 447 found by the judge ⚠️ | How much a share may be inflated |
| Interest disclosure | Not located in the document | Before the figures, alongside a dated historical brand baseline | Who benefits if the reader believes the result |
The comparison includes the campaign’s disclosed limitations. Its 29-name lens is one constraint. The five repetitions cover only 25/100 questions; the other 75 have no stability test. And evidence_excerpt is empty in 800/800 records: images exist, but there is no automatically recorded trail explaining each classification decision.
A real result, published under those rules
The hiring family contains 48 captures: 12 questions × 4 engines, n=1, in Wave 1. Under the deterministic 29-name list, the counts were Brasil GEO 19, Conversion 18, GeoStack 18, Criamente 14, SW Agência 10, Upsend 9, Bloomin 9, and Quality SMI 7. 8/48 answers contained none of the listed names.
The leaders disagree by engine, with n=12 for each: Brasil GEO, 8, on ChatGPT; Brasil GEO and Criamente tied at 7 on Claude; GeoStack, 8, on Perplexity; and Conversion, 9, on Google AI Overview.
These are absolute counts, not percentages, because of criterion 8. In the wider repeated subset, 51/100 question–engine pairs had no listed brand in a majority of the five answers. Four leaders on the same day for the same 12 questions, with more than half the repeated field unoccupied, do not establish ownership of a category.
In the historical 2026-08-06 baseline, before the current guide library, the author’s brand was absent in all 48 hiring-family captures, as in all 800 campaign captures. A brand with 15 years of off-site authority appeared in 19/100 and was absent in 81; the gap among the top three was less than 1.3 binomial standard errors, making the ranking unstable. Another brand with four months of concentrated content appeared in 15. These dated category findings do not assess today’s SWAS strategy or current SOV.
The external context, which challenges both sides
The critical survey by Olivier Martinez, submitted on July 15, 2026, examined 45 studies, from November 2023 to July 2026. None established stable, longitudinal, cross-platform causality.
SparkToro’s study by Rand Fishkin and Patrick O’Donnell, involving 600 volunteers and 2,961 runs, found high inconsistency for individual questions. Christopher Penn’s discussion of measuring AI visibility similarly challenges advice offered without a method.
The academic baseline is Aggarwal et al., KDD 2024, with GEO-bench’s 10,000 questions and 10 engines. It is a methodological reference for a commercial study, not permission to omit one.
One scope limit applies throughout this page: Google here means Google AI Overview. The Gemini app was not measured. We have no capture mechanism for it and make no claim about it.
Frequently asked questions
Is there a Brazilian study of which brands AI recommends?
There are a few, and each needs to be assessed by what it discloses. GeoStack’s Citation Share Report discloses 1,000 answers = 200 questions × 5 platforms, with one run per question per platform, but not the execution date, capture mechanism or screenshots. Our campaign discloses 800 captures, 100 questions, four engines, browser capture, 800 screenshots, timestamps for each capture and a file of verbatim questions. These are different slices of the problem. A study can be correct and still impossible to verify independently.
What does a study need to disclose to be trustworthy?
Nine things: execution date, exact denominator, capture mechanism, repetitions per question, capture evidence, verbatim questions, classification rule, measurement-lens size, and interests. These are disclosures rather than a guarantee of truth. A study meeting all nine can still be wrong; one providing only two cannot be checked. The point is to let an outside reader tell the difference.
Why do studies rank different brands first?
They use different engines, rules, repetition counts and measurement lenses. Even within our 48 hiring captures, collected on the same day for the same 12 questions, the leaders differed: Brasil GEO on ChatGPT, Brasil GEO and Criamente tied on Claude, GeoStack on Perplexity, and Conversion on Google AI Overview. Measuring one engine once is not inherently wrong; it supports a narrower claim.
Can I trust a brand’s share-of-voice percentage in AI?
Treat it cautiously unless the size of the tracked brand list is disclosed. Our list contained 29 names, while the judge found 447 in the same answers — 418 outside the list. Wyse alone appeared in 49/800 captures without belonging to our list. A fraction normalized across 29 names is larger than the equivalent fraction across 447. Ask for absolute counts and their denominators.
Is a study useful if it meets only some of the nine criteria?
Yes. Discarding every incomplete study is the opposite mistake to believing every number uncritically. The criteria limit the scope of the conclusion, rather than deciding whether the work has any value. A study that discloses its denominator and repetitions but omits the execution date may describe a real result whose current relevance cannot be checked. A denominator alone can support a hypothesis, not a comparison. GeoStack’s Citation Share Report, as read on August 6, 2026, discloses 1,000 answers = 200 questions × 5 platforms, one run each, but not execution date, capture mechanism or screenshots. That does not make it false; it makes it independently unverifiable on those points. Use it within those limits, and do not combine studies with different criteria into a single ranking chart.
Who is writing this, and the disclosure of interests
I am Mateus Gomes, operator of murmur.marketing, a SWAS combining proprietary software to measure and track SOV with specialist strategy and execution. We publish this work and sell the service it discusses. I disclose that interest while explaining how to assess studies produced by interested authors.
The campaign’s own result also needs a date and scope: in the 2026-08-06 baseline, before the current guide library, citation_kind = absent in 800/800; cold discovery was 0/360. One failed capture remained in the denominator; the lens covered 29 names out of at least 447; n=5 applied to 25/100 questions; and evidence_excerpt was empty in 800/800. These historical results are not an assessment of today’s SWAS strategy or SOV.
Conclusion
The value of a study about which brands AI recommends lies in what it discloses: execution date, denominator, mechanism, repetition, evidence, verbatim questions, classification rule, lens size, and interests. These nine criteria do not divide studies into right and wrong. They divide what an outside reader can check from what that reader must take on trust. Check those disclosures first.
Public documents disclose different parts of the list, and the campaign has its own documented limitations. None of this establishes a category owner: 51/100 repeated question–engine pairs have no listed brand holding a majority. To discuss the method or challenge a reading, contact Mateus Gomes on LinkedIn.