Translation note. This is the English version of the approved Portuguese article. The Portuguese version remains the evidence record for original AI replies. Quotes rendered in English below are identified as translations; proper names, source URLs, dates, denominators and limitations are retained.
Weekly to monitor, monthly to decide — and the right frequency is not chosen by preference; it is derived from the noise floor in your own question universe. Measuring more often than that does not produce more information: it produces more fluctuation, which gets read as movement and becomes a wrong decision.
I write as Mateus Gomes, operator of murmur.marketing, a GEO operation combining proprietary measurement software with specialist strategy and execution to grow SOV. The cadence below is grounded in a dated repeatability experiment: the same question universe was run twice without changes to estimate how much the instrument varies on its own.
Summary
- Measure noise before signal. I ran 387 questions from a real Brazilian e-commerce client universe 3 days apart, with no intervention: 98% returned the same verdict.
- In ChatGPT, 7 of 387 questions fluctuated (Δ +0.3 percentage points); in Claude, 6 of 386 (Δ 0.0). Both were pre-intervention — the delta is pure engine variance.
- The resulting scale is: movement above 3 to 4 percentage points in the overlap set is real; below that, the instrument is breathing.
- Weekly is the monitoring cadence; monthly is the decision cadence. Daily only makes sense in an incident window, and even then only with the floor known.
- One run is invalid by construction. Running n=1 and calling it a result is the most common error, and instability has already been measured by third parties in 2,961 runs.
- The real cost per run is calculable: each capture takes 17 to 51 seconds, depending on the engine. The constraint is machine time, not licensing.
Noise first, then signal — and almost nobody does it in that order
The question “how often should I measure?” sounds logistical, but it is statistical. You cannot know whether a weekly variation means anything without knowing how much the number varies when nothing changes.
That is what I did. I took a real universe of 387 questions from a Brazilian e-commerce category and ran the same measurement twice, three days apart, with no intervention between them — nothing published, changed or submitted.
| engine | questions in the overlap | first run | second run | delta | questions that fluctuated |
|---|---|---|---|---|---|
| ChatGPT | 387 | 22.2% | 22.5% | +0.3 pp | 7 |
| Claude | 386 | 21.2% | 21.2% | 0.0 pp | 6 |
About 98% of questions returned exactly the same verdict. The ones that changed did not change because anyone did anything — they changed because the engine is probabilistic.
That yields the only significance scale I consider defensible in this category: movement above 3 to 4 percentage points in the overlap set is real; below that, it is noise. It is not a convention borrowed from another discipline. It is the measured floor of the instrument that produces the number.
And that is why frequency matters: measuring faster than the real change cycle only raises the chance that you react to one of the seven questions that fluctuated on its own.
The cadence I operate, and why
| cadence | purpose | what it answers | what it does not answer |
|---|---|---|---|
| weekly | monitoring | is the trend alive? did an engine stop citing entirely? | whether a 2 pp variation means something |
| monthly | decision | did the investment change the score above the floor? | cause — that needs design, not cadence |
| per incident | diagnosis | what happened to that question, in that engine | trend |
| daily | almost nothing | — | ⚠️ produces a pretty series and a weak conclusion |
Weekly, per client, staggered. One client per weekday, not all on Monday — the constraint is machine capacity, and putting everything on one day creates a queue that delays everyone. Weekly is fast enough to notice that an engine stopped citing and slow enough not to confuse breathing with movement.
Monthly to decide. Publication takes time to become a citation, and timing is unpredictable: at one company of mine, the first citation of a /geo/ directory in ChatGPT came at T+5 days; in another case I monitor, it came at T+28 in Perplexity and T+30 in ChatGPT. With that spread, drawing an investment conclusion over a seven-day window is drawing a conclusion from nothing.
Daily, almost never. The exception is an incident window — a sudden drop, a suspicious deployment, a robots change. Outside that, a daily series in a field with a 2-to-3-point breathing floor produces a graph full of spikes that correspond to no event.
The error that comes before frequency: running n=1
Before deciding how often to run, you must decide how many times per run. Here, the error is nearly universal.
One run is invalid by construction. The same prompt, in the same engine, on the same day, returns different answers. A screenshot of one question is a sample of size 1, and its margin of error is large enough to reverse a conclusion.
This is not peculiar to my instrument. Rand Fishkin and Patrick O’Donnell’s SparkToro research — 600 volunteers, 2,961 runs, 12 categories, across three engines, fieldwork in Nov–Dec/2025 — concludes that brand recommendation is highly inconsistent between runs.
I see the same pattern in my material, clearly. The question that originated my project ran five times in each of the four engines on 2026-08-06: the most cited name on my list appeared in 5/5 in ChatGPT, 4/5 in Perplexity, 0/5 in Claude and 0/5 in Google AI Overview. Had I run it once in the wrong engine, I would have concluded the opposite of what all 20 show.
The practice: higher repetition for questions that decide something (commercial-intent questions), n=1 in the broad tail used for coverage. That is how I built my own campaign — n=1 in the 100 questions and n=5 in the 25 priority questions, totaling 800 captures.
The real cost of a run, so the calculation closes
Frequency is constrained by machine time, and in my case that calculation is public.
| engine | average time per capture |
|---|---|
| Google AI Overview | 17 s |
| Claude | 42 s |
| Perplexity | 48 s |
| ChatGPT | 51 s |
Measured across about 13,000 captures. A serial slot delivers between 1,700 and 1,900 captures per day. Do the math for your universe: 300 questions × 4 engines × n=1 equals 1,200 captures — it fits in one day on one slot. The same 300 at n=5 are 6,000, and then require parallelism or several days.
That is why cadence is not a preference: it is what fits on the machine after you choose n. And if your provider cannot tell you how long one of its runs takes, it probably is not running in a browser.
The question that changes everything: measuring by API or browser
It matters to record this because it affects frequency in a non-obvious way. API measurement is much cheaper per query, which tempts providers to increase frequency — but it measures a different number. A provider that sells API mode compared 1,000 prompts between OpenAI’s API and chatgpt.com and found 96% divergence in citation, brand order and ranking. ⚠️ The provider does not state the date on which this test ran, and I do not invent one.
If the two sources disagree on nearly everything, measuring the wrong thing ten times faster is not an improved cadence. The distinction is in GEO measurement by API or browser.
What cadence does not solve
Commercial context: Murmur provides measurement as part of a broader SWAS operation: proprietary software is paired with specialist strategy and execution. The campaign cited here was a before the current guide content was published baseline on 2026-08-06 (citation_kind = absent in 800 of 800 captures); it is not a current visibility score. The ten prompts that named the brand are kept separate from the 712 unprompted captures.
It is also worth saying whom I am measuring against, because that calibrates what cadence can detect. Under the deterministic rule over a declared list of 29 names, in the same 100 (question, engine) pairs on 2026-08-06: 51 of 100 have no owner at all. Conversion wins 19 and is absent in 81; GeoStack has 15 and Brasil GEO 14 — and all three are separated by less than 1.3 binomial standard errors, which means this is not a stable ranking. In a field like this, a weekly series showing someone change position is nearly always displaying noise, not competition. It is the same lesson as the floor, applied to reading competition: without a denominator and a margin, rank movement is decoration.
Three limitations, so the scale can be scrutinized:
1. Cadence does not produce cause. Measuring every week shows that a number changed, not why. The critical survey arXiv 2607.14035 (Olivier Martinez, 2026-07-15) reads 45 studies and concludes that none of the reviewed techniques demonstrates a stable, longitudinal, cross-platform causal effect. A dense time series does not repair the absence of design.
2. Comparing runs requires an overlap set. If the question universe changed between two measurements, the delta mixes a change in the world with a change in the scale. The comparisons above hold because they cover the same 387 and 386 questions.
3. The instrument also drifts. Improving the detector between two runs produces “improvement” that is a change in the measuring device. When I change the scale, I say so — and do not count it as a result. Christopher Penn makes this market criticism from the method side: much of the measurement advice sold here comes without one.
Frequently asked questions
How do I find the noise floor of my own universe?
Run the same measurement twice without changing anything between them, a few days apart, and count how many questions changed verdict. It costs one extra run, which pays for itself the first time it prevents a wrong decision. In the universe of 387 questions where I did this, about 2% fluctuated on their own, and that number defined the point from which I call a variation a result. Without this step, any significance threshold you adopt is borrowed from something else.
Does measuring every day not give me a richer series?
It gives you a denser, not a more informative, series, because the added density is filled by variation in the engine itself. With a two-to-three-point floor of natural fluctuation, a daily graph displays peaks and valleys that correspond to no event, and the cost is not machine time: someone will look at the peak and ask for an explanation that does not exist. The legitimate exception is an incident window, when you want fine temporal resolution to locate an event you know occurred.
Do I need to measure all four engines every time?
Yes, if comparing runs is to carry any weight, because engines disagree with each other more than one run disagrees with the next. In the question that originated my project, with five repetitions in each engine on the same day, a brand appeared in five out of five in one engine and zero out of five in another. Measuring only the engine in which you perform well and comparing it with the prior week creates a true series about a slice you chose, which is another way of choosing the result.
How many repetitions per question, then?
Use more repetitions where the answer decides money, and n=1 in the tail used for coverage. I operate n=5 for commercial-intent questions and n=1 for the rest, which in my own campaign totaled 800 captures from 100 questions. The criterion is not purely statistical; it is economic: repetition costs machine time linearly, so it belongs where a reversed verdict would change a budget decision.
How long in advance can I expect a publication to appear in measurement?
There is no reliable timeframe, and that is the honest answer. At one company of mine, the first citation appeared five days after publication; in another case I monitor, it took twenty-eight days in Perplexity and thirty in ChatGPT. With that spread, the design that works is to measure monthly for decisions and not interpret the first weeks, because in that interval absence of movement cannot distinguish “it did not work” from “it has not been crawled yet.”
Who wrote this, and the declaration of interest
Mateus Gomes operates murmur.marketing, a Brazilian GEO operation combining proprietary measurement software with specialist strategy and execution to grow SOV. The cadence recommendation comes from repeatability analysis; the 2026-08-06 brand baseline predates the current guide collection and is not a current performance score.
Conclusion
Measurement frequency is a consequence, not a choice: it comes from your universe’s noise floor, the n you chose per question and the machine time a run costs. Measure noise first, with two runs without intervention — in the universe where I did that, 98% of verdicts repeated, and the scale of “movement above three to four points” was born there. After that, weekly monitors, monthly decides, and daily is only for incidents. No cadence delivers cause: for that you need design, and the literature has not settled that part. To discuss cadence for your universe, contact Mateus Gomes on LinkedIn.