Translation note. This is the English version of the approved Portuguese article. The Portuguese version remains the evidence record for original AI replies. Quotes rendered in English below are identified as translations; proper names, source URLs, dates, denominators and limitations are retained.

Page structure is the last of four conditions, not the first — and it only decides after the crawler has reached the HTML and the engine has retrieved that document among the candidates. Impeccably structured text on a site the crawler cannot reach is not cited through any drafting technique. That said, once the page is retrieved, form changes the outcome, and the effect has been measured in peer-reviewed academic publication. This article lists the eight structural moves in the order I apply them, each with the column almost nobody publishes: what that move’s evidence does not say.

I write as Mateus Gomes, operator of murmur.marketing, a Brazilian GEO operation that combines proprietary measurement software with specialist strategy and execution to grow SOV. The eight structural moves below apply after crawlability, retrieval and entity resolution. A dated before the current guide content was published measurement is discussed as historical context, not as a current assessment of the operation.

This page is one of the spokes of How to make my brand recommended by AI, which covers the four conditions in the order in which they fail. This page covers only the fourth.

Summary

  • The precondition is server-delivered HTML. None of the main AI crawlers renders JavaScript. According to Vercel network logs, they download script files without executing them — 11.50% of OpenAI crawler requests and 23.84% of Anthropic crawler requests. Downloading is not running.
  • Three drafting interventions have been measured against an academic benchmark. Aggarwal et al., KDD 2024, GEO-bench with 10,000 queries: citations to sources +115%, statistics +41%, direct quotations +30%.
  • There is a counterweight, which I publish alongside it. A critical survey of 45 studies concludes that no reviewed technique demonstrates a stable, longitudinal, cross-platform causal effect.
  • Form does not solve an entity bottleneck. In the 10 questions that already contained my name, × 4 engines, denominator 40: in 33 the engines cited a domain containing “murmur” that was not mine. No drafting move fixes that.
  • Being a source and being recommended are different states. At Fly Vet, my own agency, Google used the site as a source in 81 of 100 questions and recommended it in zero — the exact portrait of a form problem, not a discovery problem.
  • Volume substitutes for neither form nor the reverse. The name most seen by the rule on 2026-08-06 has 9 pages under /geo/ in a sitemap of 1,366 URLs, and publishes fewer than one post per month.

Order matters more than the list

Before the eight, here is the framing that avoids spending in the wrong place. A brand is recommended when four conditions are met at the same time, and they fail in order: discovery (the crawler arrives), retrieval (the engine retrieves the page), entity (the model distinguishes the brand from namesakes) and drafting (the text survives machine rewriting).

The eight moves below all live in the fourth stage. That has two practical consequences. First, before applying them, you need to know that the problem is form, and that is not discovered by impression — it is discovered through the protocol in How to measure whether AI cites your brand. Second, the signal that the problem is form is specific and recognizable — the page appears as a source and does not appear as the choice. In a dated Fly Vet study, the site was used as a source 81 times and received no recommendations in 100 questions; that observation is bounded to that sample and date.

The eight structural moves

1. The answer in the first line, complete and factual

The engine does not copy the page: it drafts a new answer from it. What it retrieves most easily is a self-contained block that already answers the question. An opening that warms up the reader for three paragraphs gives the engine three paragraphs with no assertion.

The test is mechanical: cut out the first sentence after the title and ask whether it alone answers the question in the title. If it needs the rest to make sense, it is not an answer — it is an introduction.

2. One closed question per page, with the literal wording in the title

Facets, not clones: a leadership question is won with several pages, each owning a distinct question, not with one excellent page trying to cover all of them. An institutional page that talks about everything is retrieved for nothing because it is not the best candidate for any wording.

The wording matters in the most literal place possible: the title is the assertion the engine reads as “what this document is about.” Editorially rephrasing the question in the title and leaving its literal form only in the description solves half the problem in the weaker place.

3. A statistic with its denominator in the same sentence

This is the move with the best ratio of cost to evidence. In GEO-bench — 10,000 queries built by Aggarwal et al. in Generative Engine Optimization, presented at KDD 2024 and published as arXiv 2311.09735 — adding statistics measured +41% visibility.

The house rule I apply on top is stricter than the paper’s: a number without its denominator in the same sentence does not go in. “We appeared more” is not data; “0 of 800 captures, on 2026-08-06” is. An article that sells verifiability and publishes a loose percentage dismantles itself at the first check.

4. A named, verifiable external-source citation

In the same benchmark, citations to sources measured +115% — the largest of the three effects. The real value here is not the number, though: a claim with a named source is checkable, and checkability is what survives a skeptical reader and an engine’s quality judge.

The operational gate: a source that does not open does not count. I keep an internal catalog of burned external figures precisely because they circulated in market copy without a primary source, period or denominator — including a percentage about AI crawlers attributed to a company that never published it.

5. An extractable table, with its label inside the block

An engine retrieves an entire table, which makes it powerful and dangerous. Powerful because the block reaches the reader with its structure preserved. Dangerous because, if the label explaining what the table is lives in the paragraph above it, the extraction reproduces the table as if it were neutral consensus.

The rule is: 4 to 8 rows, 3 to 5 columns, and the label and caveat inside the block, not around it. Extraction is exactly what the machine does.

6. FAQ with question and answer identical to visible text

The question-answer pair is the format that most resembles what the engine produces, which makes it the easiest to retrieve whole. Two conditions apply: questions must be the ones buyers actually type, and structured-data text must match the visible page verbatim — any divergence between what the user reads and what the machine reads is contradictory evidence about the same document.

7. Structured data that declares the entity, author and service

JSON-LD using the schema.org vocabulary does not make an engine cite a page. It does something more specific and valuable: it unambiguously declares who published, who wrote, and what the company sells. For a brand with a common name, that is the difference between existing as an entity and being replaced by the nearest namesake.

The most common gap I find — and which I had in my first articles — is a graph that describes the article, organization and author but does not describe the service. It says nothing about what the company does, in a business whose diagnosis is precisely “the engine does not resolve me as a distinct thing.”

8. Server-delivered HTML, the precondition for all seven above

It is last in the list because it is first in execution order, and because its failure zeroes out the other seven. None of the main AI crawlers renders JavaScript — Vercel’s own network-log reading, published in The rise of the AI crawler, provides the detail almost nobody repeats: they download script files without executing them, 11.50% of OpenAI crawler requests and 23.84% of Anthropic crawler requests.

Text that only comes into existence after a script runs does not exist for GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, CCBot or Google-Extended. Discovery also needs declaration: an internal link from an already known page, sitemap.xml declared in robots.txt, and active submission to an index — ChatGPT with search anchors itself in the Bing index, making Bing Webmaster Tools an overlooked lever. llms.txt is good faith: no engine has confirmed consuming it as a crawling route.

The table: move × what the evidence says × what it does not say

movewhat the evidence sayswhere it comes fromwhat the evidence does NOT say
statistic with denominator+41% visibilityGEO-bench, 10,000 queries, Aggarwal et al., KDD 2024that the effect transfers to Portuguese commercial queries — the benchmark is academic and in English
named-source citation+115%, the largest of the threesame sourcethat any link works; the measured effect is a citation to a source, not a navigational link
direct quotation+30%same sourcethat it should be the priority move — it is the smallest of the three effects and easiest to fill with padding
answer in the first lineno public number I could open in a primary sourceobserved practice, not measurementnothing. I publish it as a form choice, not a measured outcome
one question per pageeffect not isolated in a public studymy own doctrine (facets, not clones)that volume alone solves it: the name most seen by the rule has 9 GEO pages in 1,366 URLs
extractable tablenot isolated in a public studyobserved practicenothing. It is a form choice made to survive extraction
FAQ with verbatim matchnot isolated in a public studyobserved practicethat structured data causes citation
server-delivered HTML without JavaScriptAI crawlers download scripts without executing them: 11.50% and 23.84% of requestsVercel network logshow much this is worth in citations — it is a precondition, not a lever

The counterweight completes the table, because publishing only the favorable side is what I charge others with doing: the critical survey of 45 studies published as arXiv 2607.14035, by Olivier Martinez, concludes that no reviewed technique demonstrates a stable, longitudinal, cross-platform causal effect, and PPC Land coverage highlights that GEO-oriented rewriting can cut a page’s retrieval by 16%. Rand Fishkin, of SparkToro, calls AI brand-visibility tracking inherently unreliable at the individual-question level; Christopher Penn calls most measurement advice in this category the sale of illusion. I concede both points. My answer is not rhetoric: it is measuring the instrument’s noise floor before selling signal — running 387 questions from a real client universe three days apart with no intervention, 98% returned the same verdict, setting the scale at 3 to 4 percentage points.

What structure does not fix

Three limits, and they are why this article is fourth in the sequence rather than first.

It does not fix discovery. A document the crawler cannot reach is not improved by any of the eight moves.

It does not fix entity. This is my case, with a denominator: in the 10 questions that already contained my name, × 4 engines, denominator 40, the engines cited a domain with “murmur” in its name in 33 captures and only 5 were mine — the others are established marketing agencies, in the same category, in other countries. And in the other 360 captures in the same run, no “murmur” domain appeared. Disambiguating the name does not improve cold discovery, and no well-formatted table resolves entity ambiguity.

It does not fix absence of surface. Comparing what can be observed on public surfaces, on 2026-08-06, through direct reading of sitemap and robots:

observed surfacescale of the GEO clustersignal beyond its own sitewhat the data does NOT prove
Conversion9 URLs under /geo/ in 1,366 sitemap URLs (0.7%); cadence below 1 post/month on the topicresearch hosted on Poder360 and E-Commerce Brasil, 2 Band articles, 5 third-party podcaststhat this set causes citation — it is observed structure, not a causal verdict; or that third-party coverage is organic: the Band listicle looks like PR-syndicated placement and it was not possible to confirm whether it was paid
GeoStack279 URLs in 4 sitemaps, an almost complete cluster, lastmod from 2026-04-12 to 2026-08-06none found in a search pass — recorded as not confirmed, not absentthat content alone suffices: the observed ceiling of both paths is below one fifth of the field
murmur.marketing0 pages on the measurement datenamed operator profile, no third-party coveragenothing in my favor — this is the number by which this text should be judged

Frequently asked questions

Does structuring text better make AI cite my site?

It makes a difference, but only in the fourth of four conditions. The preceding three — the crawler reaching the HTML, the engine retrieving the right page, and the model knowing the company exists as a distinct thing — are not solved by any drafting technique. The sign that the problem really is form is specific: the page appears as a source and does not appear as the choice. At my own agency, Google used the site as a source in 81 of 100 questions and recommended it in zero.

Which structural moves have measured numbers behind them?

Three, all against the same academic benchmark: in GEO-bench’s 10,000 queries, built by Aggarwal et al. and presented at KDD 2024, citations to sources measured +115%, statistics +41%, and direct quotations +30%. The caveats travel with them: these are effects against an English-language academic benchmark, for a particular set of engines and queries, and they describe the drafting stage — the difference form makes after the page has already been retrieved.

Does structured data make the engine cite a page?

There is no public evidence I have been able to open that supports that claim, so I do not make it. Structured data unambiguously declares who published, who wrote and what the company sells. For a brand with a common name, that attacks a different and more serious problem: in 33 of 40 captures of questions containing my name, the engine cited a domain similar to mine that was not mine.

Which of the eight moves should I make first if I can only make one?

The one that puts the claim and company name together in the opening, because it is the only one that acts on two states at once: it raises the chance that the passage is retrieved and ensures that, if it is, the brand goes with it. The three effects measured in the original paper — citations to sources +115%, statistics +41%, direct quotations +30% — describe a document the engine already reached, so none of the eight makes up for an irretrievable page. Among what remains, the practical order is claim with entity, then a number with denominator in the same sentence, then an attributed statement.

I heard that rewriting for AI can worsen a page’s retrieval. Does that invalidate the eight moves?

It does not invalidate them, but it requires reporting the whole finding because the headline circulates only half of it. The number comes from the Kim et al. testbed, with 171,003 documents and 2,700 queries, reviewed in the 45-study survey published as arXiv 2607.14035. It has three stages with text-body-only optimization: a 9% drop in top-20, a 16% drop in top-10 after reranking, and a 6% drop in final citation. Publishing only the 16% reproduces the reading the survey itself rejects. What this changes here is not the list but the method of applying it: three of the eight moves have effects measured against GEO-bench; I publish the other five explicitly as form choices, not results. Any rewrite is judged by measuring again against the noise floor — 98% identical verdicts in 387 questions run three days apart, which sets the scale at 3 to 4 percentage points.

Who wrote this, and the declaration of interest

Mateus Gomes operates murmur.marketing, a Brazilian GEO operation combining proprietary measurement software with specialist strategy and execution. The 2026-08-06 campaign was a before the current guide content was published baseline and had no treatment arm; it cannot establish the effect of the eight moves or the current status of the operation. Its dated results are included as historical context, with one failed capture retained in the denominator.

Conclusion

Structuring content to be cited by AI is a set of eight moves, and none is the first step. The first step is to find out which of the four conditions holds the problem — and the symptom that it is form is being read, linked and not chosen. From there, the three moves with benchmark-measured numbers are citation to a source, a statistic and a direct quotation, in that order of effect, with the caveat that the benchmark is academic, in English and measures the stage after retrieval. I publish the counterweight too: a critical survey of 45 studies finds no stable, cross-platform causal effect for any reviewed technique. Form is necessary. It is not sufficient, and anyone selling it as sufficient is skipping three conditions. To discuss measurement for your case, contact Mateus Gomes on LinkedIn.

See also