A question set is a file, not a list of ideas: each row contains the verbatim question, its role, the reason it exists — written before the campaign — and, afterward, the result. Without a declared role per row, figures from different groups end up combined. That combination inflates precisely the metric you care about. It is the costliest and most common error in AI brand measurement, and it is not a capture error: it is an instrument-design error.
I am Mateus Gomes, operator of murmur.marketing, a Brazilian GEO operation combining proprietary measurement software with specialist strategy and execution to grow SOV. The question set should support an ongoing cycle of diagnosis and action, not a one-off score. The dated campaign below predates the current guide collection and is included as a historical method example.
This page builds on How to measure whether AI cites your brand, where the full protocol is described. Here we address only its first decision: which questions belong in the set.
Summary
- The question set is the instrument. If it changes between two rounds, there is no comparison: there are two snapshots of different things.
- Eight columns per row, the most important being
rationale: why this question was included, written before running it. - Four roles, designed to prevent inappropriate aggregation. My set has 34
atacarquestions (compete), 26rastrear(track), 30baseline, and 10defender(defend). These are the original role identifiers. - The trap: a prompt that already contains the brand measures recognition, not discovery. In my campaign, all 88 hits in 800 captures came entirely from the ten questions already naming it. Outside them: 0 of 712.
- The denominator belongs to each group, never just the aggregate:
atacar0/136,rastrear0/104,baseline0/120, and the group already naming the brand, 40/40 of pure echo. - Retaining the verbatim questions in version control and making them available on request turns a report into something a third party can repeat.
1. What a question set is, and what it is not
A question set is not a keyword list. A keyword is a fragment. People type full sentences into generative search, often in the first person and with context. For example, “minha clínica não aparece no Google, o que eu faço” (English translation: “my clinic does not appear on Google, what should I do?”) is a question, not a term.
Three practical requirements, all inexpensive:
- Fixed and versioned. A dated file in a repository. The next comparison runs the same rows. A set that “improves” every round destroys the time series.
- The reason for each row written beside it, beforehand. Months later, nobody remembers why a question was included, and without that record the figure means nothing.
- Split by role. Section 3 covers this: roles let the instrument measure more than one thing without confusing them.
2. Eight columns for each row
The set I ran on 2026-08-06 has exactly eight fields per row, and none is decorative:
query_id— a stable identifier. It reconnects a capture, a verdict and a screenshot months later.query_run— the verbatim question, exactly as it will be typed. No normalization, accent corrections or improvements to the buyer’s grammar.tipo— the topic group. My set has ten:ranking_agencia,ranking_ferramenta,comparativo_conc,metodo,taxonomia,prova,category_head,ceticismo,personaandbranded. These original identifiers represent agency rankings, tool rankings, competitor comparisons, method, taxonomy, evidence, broad category questions, skepticism, persona and branded questions.chance— competitive priority, fromp1top3. It determines where expensive repetitions will be spent.role— the measurement role:atacar,rastrear,baseline,defender. This is the column that prevents inappropriate aggregation; section 3 is devoted to it.gap_hypothesis— the hypothesis about what is missing if the brand does not appear: content, entity recognition, brand or sentiment. Written beforehand, it becomes falsifiable.segmento— the audience or industry segment to which the question belongs.rationale— the sentence explaining why the row exists. After the campaign, each row also has a result. Together, they become two things: a content brief and a priority order.
rationale is the column almost nobody writes, and it supports everything else. A real example
from the first row of my set: “qual a melhor empresa pra fazer GEO no Brasil” (English
translation: “what is the best company for GEO in Brazil”) was included because it was
the exact question in the screenshot that started the business, and would serve as a
reference: if its answer changes one day, the project worked. The measured question was in
Portuguese. This is a success criterion declared before the result, the opposite of choosing
a metric after seeing the number.
rationale has a second use that connects measurement to production: every row where the brand
is absent is a brief for a page. What to put on that page belongs to another topic, but the
reference experimental work on what changes a generated answer is
Aggarwal et al., KDD 2024. It measured visibility gains of
+41% with statistics and +30% with direct quotations on a 10,000-query benchmark, pointing
toward pages with numbers and sources rather than adjectives.
3. Four roles: the column that prevents inappropriate aggregation
Different roles measure different mechanisms and cannot be combined into a single figure. Here is the breakdown of the set I ran, with each group’s denominator in the first wave of the 2026-08-06 campaign: one repetition per question across four engines.
| Role | Questions | Captures | What it measures | What it does not support concluding |
|---|---|---|---|---|
defender — defend: the brand is in the prompt | 10 | 40 | Whether the engine recognizes a name already supplied | Nothing about discovery. The result was 40 of 40, pure echo rather than visibility |
atacar — compete: method, taxonomy and evidence | 34 | 136 | Whether the brand appears where it should be the answer | Nothing about market demand. Result: 0 of 136 |
rastrear — track: who the competitors are | 26 | 104 | The competitive landscape described by the engines | Nothing about the brand itself. Result: 0 of 104 |
baseline — broad demand | 30 | 120 | Whether the category already has an associated vendor | Nothing about purchase intent. Result: 0 of 120 |
The four rows add up to 400 captures, the first-wave total. You can verify the sum of the denominators: that is how you know no question was counted twice or forgotten.
Notice what this breakdown reveals without further work. The dated baseline is not uniform by accident: it appears in each group with its own denominator. One aggregate would hide that and conceal the only group with hits, which is precisely the one that measures no visibility at all.
4. The central trap: a question that already names the brand
This is the line between an honest figure and an inflated one. Recognizing it in my own report was a costly lesson.
The deterministic rule recorded 88 “murmur” hits in 800 captures. A naive reading would publish “11% presence.” All 88 came from the ten questions that already contained the name: the engine was echoing the question before saying it had not found the brand. Outside those ten questions: 0 of 712. The honest figure was zero.
A second finding in that same group was possible only because the role was declared. Among the 40 captures of questions naming the brand, 33 cited a domain with “murmur” in its name, and only 5 were mine. Across the other 360 captures of unprompted discovery, no “murmur” domain appeared, neither mine nor a namesake’s. These are two different problems: name ambiguity and absence from the category. Combining the groups into one figure would merge the problems, even though fixing one does not fix the other.
Operational rule: include brand-naming questions because they measure something real, but give them their own role and denominator. Never combine them with unprompted discovery.
5. Where questions come from
Four sources, in order of value:
- The buyer’s exact wording. In the Brazilian market, people type “aparecer no ChatGPT” (English translation: “appear in ChatGPT”), not the category acronym. A question set written in the vendor’s vocabulary measures the vendor.
- The question that started the problem. Almost every project starts with a screenshot of someone asking something specific. Include that question verbatim as a success criterion.
- The deeper follow-up questions a buyer asks before buying: method, evidence, skepticism and comparison. These are the questions nobody puts in the set, and where the brand often disappears.
- The competitive landscape, represented by
rastrear: questions about others, whose answer maps the field rather than the brand.
6. What each group returns in practice
It helps to see the kind of answer each role produces, because it changes the interpretation.
The hiring group — 12 questions × 4 engines = 48 captures, using the deterministic rule on a declared list of 29 names on 2026-08-06 — returned these absolute counts: Brasil GEO 19 · Conversion 18 · GeoStack 18 · Criamente 14 · SW Agência 10 · Upsend 9 · Bloomin 9 · Quality SMI 7. Two caveats belong with these figures. The four engines disagree about the leading name: ChatGPT names Brasil GEO in 8 of 12; Claude names Brasil GEO in 7 and Criamente in 7; Perplexity names GeoStack in 8; and Google AI Overview names Conversion in 9. Also, none of the 29 names appears in 8 of the 48 captures. These are absolute counts and never become percentages: the detection list recognizes 29 names out of at least 447 extracted by the judge from the same captures.
This disagreement between engines is not peculiar to my question set. SparkToro’s research, by Rand Fishkin and Patrick O’Donnell, with 600 volunteers and 2,961 runs in ChatGPT, Claude and AI Overview across 12 categories, measured the same instability at a much larger scale. That is exactly why each row needs its role and denominator: without them, instability becomes noise with no identifiable source.
The broad-demand group returns something else. For a broad question — what GEO is, whether it is worth investing in — the engine answers with a concept, not a vendor. These are English translations of the Portuguese question topics. The most frequently cited name on the list appears 13 times in 120 captures. Across the first wave as a whole, none of the 29 names appears in 216 of 400 captures, with a median of zero distinct brands per capture. One group measures the field; the other measures whether the field exists. Combining them erases both readings.
7. An example showing that the breakdown is the finding
The best argument for dividing by role is not my zero. It is a case where a brand has results and still loses an entire group.
For Fly Vet, a company I own, the June 2026 campaign ran 100 questions per engine, one
repetition each, using the citation_kind rule. ChatGPT recommends it in 30 of 100, Claude
in 41 of 100, and Google never recommends it, using it as a source in 81 of 100.
Respectable figures.
Then look at one group. In the subset where a veterinarian describes their problem in the first person — 40 questions × 2 repetitions = 80 records per engine — Fly Vet appears 2 times in 80 in ChatGPT and zero in 80 in Claude. I win the comparison and lose the problem-led question. It is the costliest gap in my own funnel, hidden inside a healthy aggregate. Only the breakdown by question group exposed it.
The contrast also warns against promised deadlines. For its sister company Fly Med, the first citation of a directory page appeared in ChatGPT at T+5 days after publication, with a screenshot in the repository. In the same probe on the same day, Fly Vet scored zero. The same methodology and date produced opposite results.
8. How many questions, and how to version them
Size belongs to another page in this topic group, How many questions you need for reliable AI measurement, because it depends on repetitions and the capture budget, rather than preference. What belongs here is the discipline of maintaining the file:
- One version per date. The 2026-08-06 set is
universe-v1, with a hundred rows. The next round runs those same hundred before adding anything. - An addition is an addition, not an edit. Changing an existing question’s text breaks the
entire series for that
query_id. Create a new row and retire the old one with a date instead. - Retain the verbatim wording and provide it on request. The hundred questions are versioned in a file exactly as typed, and I provide them to anyone who asks. That allows a third party to repeat the experiment and separates a report that can be checked from one that must be believed.
Credit is due to others publishing in this category. The Portuguese-language AI citation study
closest to mine, GeoStack’s Citation Share Report — Bancos Digitais no Brasil (“Digital Banks
in Brazil,” translated title), which I read on 2026-08-06, declares its design in its own text:
1,000 observations, 200 prompts × 5 platforms, each prompt run once per platform. That is more
than many visibility reports disclose. It does not provide the execution date — only publication
on 2026-04-29 — the capture mechanism or the verbatim questions. This is not a judgment of the
study’s quality. It lists what each document declares, and a study can be correct without
being verifiable.
9. What a question set does not measure
Three limits, stated so readers can hold the method to them:
A question set measures answers, not demand. It does not tell you how many people ask that question. Search volume and AI behavior require different instruments. Confusing them is how a report ends up promising traffic on the basis of citations.
A question set does not establish causation. It describes the state of the field on one day. What makes a brand improve is a separate question, and the literature has not settled it. The critical survey arXiv 2607.14035 (Olivier Martinez, 2026-07-15) reviews 45 studies and concludes that none of the reviewed techniques demonstrates a stable, longitudinal, cross-platform causal effect.
The detection list limits what the set can see. Mine covered 29 names; the judge extracted at least 447 distinct brands from the same 800 captures. The most frequently cited unlisted name appeared 49 times in 800, and I did not know it existed. A well-designed question set with a poor detection list returns an incomplete map with unwarranted confidence.
Frequently asked questions
How many questions does a set need?
It depends on what you want to detect. Repetition contributes more to reliability than the number of questions. My 2026-08-06 design ran 100 questions once across four engines, with 25 priority questions reaching five repetitions in those same engines: 800 captures in total. A hundred questions with five runs for the priorities detects more than a thousand questions run once.
Should I include questions that already name my brand?
Yes, but with their own role and denominator, never combined with the others. They measure recognition, not discovery. In my campaign, all 88 deterministic-rule hits in 800 captures came from the ten questions already containing the name. Outside those ten, the result was 0 of 712, and the report’s honest figure was zero.
What should I write in each question’s rationale column?
The reason that row exists, written before running the campaign: which decision its result will inform, and what it would mean for the brand to appear or not appear there. Written beforehand, it makes the question falsifiable and becomes a content brief afterward. Written afterward, it becomes a rationalization of whatever number appeared.
Can I use the same questions for every engine?
Yes. If the goal is to compare engines, you must. Using the same wording in all four reveals their disagreement. In my campaign’s hiring group, the four engines did not agree on the leading name across the same 12 questions, and none of the names on the declared list appeared in 8 of 48 captures.
Do the four roles need an even split?
No. The split follows the decision each group will inform, not symmetry. In my 2026-08-06 set, the hundred rows comprised 34 atacar questions, 26 rastrear, 30 baseline and 10 defender. The group already naming the brand is deliberately the smallest because it measures recognition, not discovery. Every role carrying its own denominator is non-negotiable: first-wave results were 0 of 136 for atacar, 0 of 104 for rastrear, 0 of 120 for baseline, and 40 of 40 for the group already naming me. Those four parts sum to 400 captures. If a group is too small to support an interpretation on its own, that is where it shows up, not in the total.
Who wrote this, and disclosure of interest
Mateus Gomes operates murmur.marketing, a Brazilian GEO operation combining proprietary measurement software with specialist strategy and execution to grow SOV. This question-set method connects a baseline to an ongoing strategy rather than treating measurement as the outcome.
The 2026-08-06 campaign is a historical example from before the current guide collection: citation_kind = absent in 800 captures across four engines. The correction separating brand-naming prompts from unprompted discovery illustrates why the question universe and denominator must be explicit. This result is not a current visibility score. The full measurement, including its methodology and what was not measured, is published in the
Observatório GEO Brasil.
Conclusion
Building a question set means maintaining a measurement file, not brainstorming: eight columns per row, the verbatim question, a declared role, a reason written beforehand and a result recorded afterward. The role carries the weight. Without it, questions already naming the brand contaminate the group measuring discovery, and the report publishes a figure its own author cannot defend. In my case, the difference was “11%” versus zero. Retaining the verbatim questions completes the process, enabling a third party to repeat the experiment independently of whoever measured it. To discuss a question set for your situation, contact Mateus Gomes on LinkedIn.