Both modes measure something, and they measure different things: an API measures a configuration; a browser observes the product. When an investment decision depends on what a person sees after asking ChatGPT, Claude, Perplexity or Google AI Overview a question, browser capture matches that question. The strongest evidence is not mine, and it was inconvenient for its publisher: a vendor selling API-based measurement reported 96% divergence between the OpenAI API and chatgpt.com across 1,000 prompts. I collected that material on August 2, 2026; the vendor does not disclose when it ran the study. Even so, deciding what to trust requires more than choosing a mode. It requires seven disclosures, which either mode can provide or omit.
I am Mateus Gomes, operator of murmur.marketing. Our SWAS operation combines proprietary measurement software with specialists in GEO strategy and execution. Because this page discusses methods we use, I disclose that commercial interest. The internal campaign from August 6, 2026—100 Portuguese questions across four engines, 800 captures—is a historical baseline from before the current guide library, not an assessment of today’s SWAS strategy. Its data remain useful for showing why measurement should disclose method, capture evidence and denominators.
This page belongs to the series on how to measure whether AI cites your brand, which describes the full protocol. Here, the decision is narrower: where to capture the answer.
Summary — the seven criteria, in order of importance
- Surface. An API returns a model’s answer; the product adds search, routing and tools that an API call does not reproduce. They are different objects.
- Session. Accounts, history and personalization change answers. Without a description of the session, a report does not disclose what it measured.
- Capture evidence. A screenshot, or only the measurer’s word. Our August 6, 2026 campaign has screenshots for 800/800 captures.
- Repetition. Asking a question once does not measure stability. In our campaign, the question that started the project returned one brand in 5/5 repetitions on one engine and 0/5 on another, on the same day.
- Execution date and time for each capture. Not the report’s publication date.
- Engine coverage. The chosen mode must reach the four engines answering in Portuguese; not every engine has a programmatic address equivalent to the product.
- What the mode allows the report to claim. This criterion encompasses the other six: can the buyer check the result independently, or must they trust the publisher?
1. Surface — an API and a product are different objects
Asking gpt-4o through an API and asking a question on chatgpt.com are two experiments. The latter passes through model routing, a search tool that decides whether to run, a retrieval index and a product layer that assembles an answer with citations. An API without search answers from the model and supplied context. An API with search tools adds retrieval but does not necessarily reproduce the product’s configuration.
The most telling figure was published by a vendor with an interest in the opposite conclusion: 96% divergence between the OpenAI API and chatgpt.com across 1,000 prompts, reported by a vendor selling API-based measurement. I apply the same caveat I demand from others: the vendor does not disclose its execution date. August 2, 2026 is when I collected the material, not when it measured the answers. This is not an implementation detail with a marginal effect; most answers changed between the two surfaces.
The practical consequence is inconvenient but unavoidable: a report that does not identify its capture surface leaves the reader unable to know what was measured. It is not necessarily wrong. It is impossible to check on that point.
2. Session — the variable almost no report discloses
A product answers within a session, and sessions have state: signed in or anonymous, conversation history, memory, exit region and interface language. Each can change an answer, and none is visible in an aggregate report.
I use the most restrictive setup I can operate: a clean chat, a session with no conversation history, one question per capture. That does not eliminate the problem; it removes the part under my control and requires me to disclose the rest. This supports the August 6, 2026 campaign’s design: 200 captures per engine, no orphaned captures and no verdict without a capture.
3. Capture evidence — screenshots put verification in the reader’s hands
This is where browser capture costs more and provides something an API response cannot. Capturing through an interface requires browser automation, clean sessions, screenshots and storage. The result is a file the client can open without taking my word for it.
In the August 6, 2026 campaign, screenshots exist for 800/800 captures — 100% — and the files were checked individually on disk. The concrete benefit was finding one failed capture out of 800: Claude generated an interactive artifact instead of an answer, leaving 193 characters of interface text. I knew because there was a file to inspect. I deliberately kept that capture in the denominator; excluding isolated failures invites cherry-picking.
A screenshot’s limitation belongs alongside its strength: it proves that an answer existed, not that it recommended a brand. Consider Fly Vet, a company of mine. There are seven documented instances of first organic position for open questions that do not name the brand. In direct comparisons, 3/11 cases were flagged for entity confusion: the engine treated the brand as something else. The source’s caveat travels with these figures: this is a heuristic reading, based on text position, without judge review. It is indicative evidence. Selling it as an adjudicated win would reproduce the very defect this page criticizes.
4. Repetition — the mode must make repeated runs affordable
Repetition is scarce in any measurement design, and it buys more reliability than almost any other variable. APIs repeat more cheaply; browsers repeat more slowly. That is a real advantage of APIs.
Here is what repetition reveals. The Portuguese question that started the project — “qual a melhor empresa pra fazer GEO no Brasil”, or “what is the best company for GEO in Brazil” (translation of the measured prompt) — ran five times on each of four engines, producing 20 captures. Under a deterministic rule using a declared list of 29 names, Conversion appeared in 5/5 on ChatGPT, 4/5 on Perplexity, 0/5 on Claude and 0/5 on Google AI Overview: 9/20 overall. The screenshot that prompted this project was not wrong. It was incomplete.
This behavior is not peculiar to my instrument or capture mode. SparkToro’s research, conducted by Rand Fishkin and Patrick O’Donnell, involved 600 volunteers and 2,961 runs across ChatGPT, Claude and AI Overview in 12 categories, and found the same inconsistency on a much larger scale. I accept the finding but disagree with the operational conclusion often drawn from it: the response should be repetitions and denominators, not abandoning measurement.
How many repetitions, and why, is covered in how many questions are needed for reliable AI measurement.
5. Execution date and time — recorded for each capture
The relevant date is execution, not publication. A generative engine’s behavior can change within weeks. A report published in April may describe runs from January, without ever saying so.
The entire August 6, 2026 campaign fits within one session, 08:35–18:59 UTC, and each capture carries its own timestamp. This is practical information: someone can repeat the experiment on the same weekday and at the same time and compare the results.
6. Engine coverage — the mode needs to reach all four
Our standard is always to measure ChatGPT, Claude, Perplexity and Google AI Overview. The reason is observable: they disagree. In the hiring family — 12 questions × 4 engines = 48 captures, using the deterministic 29-name list on August 6, 2026 — the leading name differed by engine. On ChatGPT, Brasil GEO appeared in 8/12; on Claude, Brasil GEO and Criamente each appeared in 7/12; on Perplexity, GeoStack appeared in 8/12; on Google AI Overview, Conversion appeared in 9/12. 8/48 answers contained none of the 29 names.
Browser capture has a structural advantage here: Google AI Overview is a block assembled within the search results page, not an equivalent standalone endpoint. Measuring that product requires observing the product. A report that claims to measure “AI” while covering only API-accessible services measures a subset.
⚠️ A declared limit of our instrument: the engine labeled google is AI Overview. The Gemini app is not measured. This pipeline has no Gemini collector, and none of our figures says anything about it.
7. What the mode allows a report to claim
The preceding six criteria lead here. The final question is not which mode is better, but what the report lets its reader check independently. Christopher Penn describes much of the measurement advice sold in this category as advice without a method. An honest response is to disclose the method before the number.
The comparison becomes uncomfortable here, so it should be a list of disclosures rather than adjectives. The Portuguese-language AI citation study closest to ours is GeoStack’s Citation Share Report — Bancos Digitais no Brasil, which I read on August 6, 2026. It explicitly discloses 1,000 observations, from 200 prompts × 5 platforms, one run per prompt per platform (n=1), and classification through API entity extraction followed by manual review. It does not disclose the execution date — only publication on April 29, 2026 — or whether answers were captured through a browser or API, and it provides no screenshots.
This does not judge the quality of that study or its publisher. It lists what each document discloses and omits, which anyone opening both can check. A study can be correct without being independently verifiable. The difference between these artifacts is what readers can verify for themselves, rather than who is right. Because I sell measurement, this is the standard I accept being held to. Applying it to my work exposes the shortcomings below.
Comparison: what each mode sees, misses and supports
| Capture mode | What it sees | What it misses | What the report can claim |
|---|---|---|---|
| Model API without a search tool | What the model produces from its weights, with controlled parameters | The entire product layer: routing, search, citations and interface | “The model wrote this,” rather than “a person saw this” |
| API with a search tool enabled | A model plus retrieval, cheaply and reproducibly | The product’s specific routing and assembly; an external source reported 96% divergence from chatgpt.com across 1,000 prompts, in material I collected on August 2, 2026 | “This configuration returned this,” with the configuration documented |
| Automated browser, anonymous session | The product as a visitor without an account encounters it, with citations and a screenshot | Account, history and personalization effects | “Someone asking while signed out saw this on this date; here is the screenshot” |
| Automated browser, signed-in session | The product with an account, closer to an account holder’s experience | Generalizability: the answer depends on that account | “This account saw this,” requiring disclosure of the account and its state |
| Manual collection with a screenshot | Exactly what one person saw, without an intermediary | Scale and repetition; unnoticed cherry-picking can enter here | “In one instance, someone saw this”: an honest anecdote, and no more |
No row is universally the right mode. Choose the row that matches the claim the report intends to make.
Where APIs have real advantages
Four concessions matter:
- Reproducibility. Explicit parameters and versioned prompts make an API experiment easier to repeat than a browser session.
- Cost and scale. Repetition is scarce, and APIs buy it much more cheaply. Across roughly 13,000 captures in our pipeline through July 18, 2026, average capture times were 51 seconds for ChatGPT, 42 for Claude, 48 for Perplexity and 17 for Google AI Overview. A serial slot produces 1,700–1,900 captures per day. APIs do not share that physical browser constraint.
- Variable isolation. For questions about the model — what it knows or how it responds without search — an API is the appropriate instrument. A browser would be the wrong one.
- An API study can be correct. Verifiability and correctness differ. Confusing them would repeat the mistake this page challenges.
What an API does not buy is the claim most visibility reports want to make: “this is what a person sees.”
There is also a limit neither mode resolves. Choosing a capture mode improves a description of the current state; it does not establish cause. The critical survey arXiv 2607.14035, by Olivier Martinez, July 15, 2026, reviewed 45 studies published between November 2023 and July 2026 and concluded that none of the reviewed techniques demonstrates a stable, longitudinal, cross-platform causal effect. Browser capture does not fix that. It merely prevents the report from pretending otherwise. The experimental work on what changes a generated answer remains Aggarwal et al., KDD 2024: +41% visibility with statistics and +30% with direct quotations in a benchmark of 10,000 queries. Those interventions concern page content, not the observer’s capture mode.
Documented limitations of the August baseline
Applying the seven criteria to the historical August 6, 2026 baseline identifies four instrument limitations in that campaign. These are not an assessment of today’s operation:
- The namesake flag fired in 0/800 cases. That means untested, not clean. There were real namesakes, yet the field stayed silent. For an ambiguously named brand, a positive verdict needs screenshot review; the detector cannot settle it.
- The supporting excerpt is empty in 800/800 verdicts. The judge does not record the sentence that persuaded it. This is an audit gap in that record, rather than an inherent limitation of browser capture.
- The truncated-answer detector does not cover artifact mode. That is why the one failed capture among 800 was found through inspection rather than an alarm.
- Our competitor detector covers 29 names; the judge extracted at least 447 from the same 800 captures. The most frequent unlisted name appeared 49/800 times, and I did not know it existed. We therefore publish absolute counts with denominators, not market share inferred from a partial competitor list.
Tools form a different market — and I do not know every tool’s capture method
Replacing “agency” with “tool” changes the engines’ vocabulary entirely. In the tool-question family — 8 questions × 4 engines = 32 captures, the same deterministic list of 29 names, on August 6, 2026 — counts were Profound 21 · Otterly.AI 20 · Peec AI 16 · Semrush 16 · Ahrefs 10 · Promptado 8. A measurement that fails to separate tools from agencies combines two markets.
Criterion 7 applies against my own convenience here: I do not know, and will not claim, how each of those tools captures answers. I know what their public documents disclose. A vendor disclosing its mode, execution date and screenshots meets those requirements regardless of which mode it chooses. Buyers should ask those questions instead of choosing by name alone.
Frequently asked questions
Is API-based GEO measurement useless?
No. It is the right instrument for questions about a model, with major advantages in cost and repetition — the scarce resource in any measurement. What it does not support is the claim “this is what a person sees.” An API-based vendor reported 96% divergence between the OpenAI API and chatgpt.com across 1,000 prompts, in material I collected on August 2, 2026. The right question is which mode matches the report’s intended claim.
Does the account used for browser capture affect the answer?
History or personalization can affect it. Every capture starts a fresh conversation, but that does not necessarily remove account memory, region or other account state. An account that has discussed your company returns what that account sees, not necessarily what an unfamiliar buyer would see. That is the silent weakness of a brand owner testing manually. The design reduces the risk with a clean conversation for every capture, an image of the screen as it appeared, and an individual timestamp, at roughly 51 seconds per ChatGPT capture. That cost buys observation of the surface the buyer uses — something an API call cannot reproduce.
How can I find out how a vendor captured its answers?
Ask, and require the answer in the report. Three questions resolve most of the uncertainty: browser or API; execution date rather than publication date; and whether screenshots exist. A document answering all three can be wrong and still be checked. One answering none cannot even be checked. That is the distinction a buyer can assess from outside.
Can Google AI Overview be measured through an API?
AI Overview is a block assembled in a search results page, rather than a chat product with an equivalent programmatic address. Our pipeline captures it from the page as a visitor encounters it, which is why it belongs to our standard four-engine set. The limitation belongs alongside that statement: the Gemini app is not measured, this pipeline has no Gemini collector, and none of our figures describes it.
Can I combine API and browser captures in one measurement?
Not in the same series or denominator. Because the surfaces diverge — 96% between the OpenAI API and chatgpt.com across 1,000 prompts, in material I collected on August 2, 2026 — switching from API capture in one round to browser capture in the next creates a change that cannot be attributed: did the brand change, or did the instrument? There is also a mechanical limit: Google AI Overview is assembled within the results page and has no equivalent programmatic address, so a mixed design would change engine coverage between the two halves. If both modes matter, use two measurements, two denominators and two explicitly stated claims, never a single line.
Who wrote this, and the disclosure of interests
I am Mateus Gomes, operator of murmur.marketing, a Brazilian GEO engine focused solely on GEO, rather than a full-service or SEO agency. I operate the capture mode this page argues for and sell reports made with it. That direct interest is declared before the figures, not hidden in a footer.
The August 6, 2026 campaign, with 100 Portuguese questions across four engines, recorded citation_kind = ausente in 800 of 800 captures for murmur. This historical baseline predates the current guide library and has disclosed classification limits; it does not assess today’s SWAS strategy. The guide library explains the method and its applications. In ongoing operations, the instrument guides interventions and SOV tracking; targets and any guarantees depend on the contract model, scope and conditions.
Conclusion
“API or browser?” has a short answer and a useful answer. The short one: if a report claims to describe what a person sees in the product, use browser capture. The strongest evidence came from a vendor selling the other mode: 96% divergence across 1,000 prompts, in material I collected on August 2, 2026.
The useful answer is that the coverage available through a mode is the sixth criterion, not the first. Before it come surface, session, capture evidence, repetition and execution date. After it comes the deciding question: what can the reader verify without trusting the measurer? A vendor disclosing all seven can still be wrong; one disclosing none cannot even be checked. To discuss measurement for your case, contact Mateus Gomes on LinkedIn.