Research measures content interventions in generative systems, but results depend on the design, engine and stage being evaluated. Ask which intervention produced which effect, on which metric and under what conditions. That makes research useful without turning an experimental percentage into a universal promise.

For a business, literature supplies hypotheses. A GEO program must turn them into market-appropriate actions and monitor presence, recommendation and SOV.

The paper introducing GEO

Generative Engine Optimization by Aggarwal et al., presented at KDD 2024, introduced GEO-bench with 10,000 queries. It evaluates interventions such as adding sources, statistics and quotations, observing changes in visibility within generated answers.

Read the metric and conditions alongside the effect. Better use of a document made available to a system does not establish that a new webpage will be discovered, chosen by every engine or generate revenue.

The operational lesson is to test clarity, attribution and verifiable information when they help answer the question. Adding numbers or quotations to satisfy a formula does not necessarily reproduce an experiment.

What a literature review adds

Olivier Martinez’s review arXiv 2607.14035, submitted on July 15, 2026, examines 45 studies from November 2023 to July 2026. It discusses limits to claims of stable causal effects across platforms and over time.

The review helps distinguish post-retrieval outcomes from tests including document selection. Practically, choose an intervention with an explicit hypothesis and verify its effect in the buyer-relevant experience.

In SAGEO Arena, described in the review and attributed to Kim et al. (2026), the dataset includes 171,003 documents and 2,700 queries. For body-only text optimization, the synthesis records −9% in the top 20, −16% in the top 10 after reranking and −6% in final citations. These are findings from that test and intervention. They do not show that all content improvement harms outcomes; they illustrate why retrieval and final citation should be tracked together.

Different studies answer different questions

EvidenceQuestion it can addressWhat should not be inferred automatically
Supplied-content benchmarkDoes document presentation affect answer use?Discovery of the document on the web
Retrieval-and-generation testDoes the intervention improve the measured chain?Equal effects in every commercial engine
Repeated real-product observationsHow often does the answer change?The cause of an editorial change
Before-and-after caseDid observed outcomes improve after the work?That no other variable contributed
Citation researchWhich sources appear in the sample?Revenue or financial returns

The SparkToro study, with 600 volunteers, 2,961 executions and 12 categories, informs variability assessment. It is not a treatment arm testing a particular GEO program.

Where murmur’s data fits

The August 6, 2026 Brazilian campaign included 800 browser captures, covering 100 Portuguese questions and four engines, with five total executions for 25 priority questions. Google means AI Overview; Gemini was excluded.

It preceded publication of this article collection and forms an observational baseline. Without a treatment arm, it does not measure the effect of publishing these guides. It documents presence and informs question selection, classification and repetition. Intervention cases require their own dates, scope and evidence.

Turning research into execution

State the hypothesis before making a change: “the page does not address implementation requirements; we will add criteria and documentation, then observe these questions”. Define the starting version, action, common universe and review window.

Monitor access, visible sources and recommendation without substituting one metric for another. Where feasible, retain comparable groups and avoid simultaneous changes that make interpretation impossible. Findings that do not confirm the hypothesis inform the next cycle.

A contractual SOV-growth guarantee is a commercial commitment with defined conditions. It does not turn research into proof of fixed rankings or guaranteed sales.

Frequently asked questions

Does research prove that GEO increases sales?

The studies discussed here mainly address retrieval, visibility and answers. Sales require commercial monitoring and a specific attribution design.

Can a paper’s percentage gain become my target?

Use it to understand a hypothesis, not as an automatic forecast. Set targets from your baseline and project conditions.

Why compare before and after without isolating every cause?

It documents progress and informs decisions. Recording changes, preserving the universe and adding feasible controls strengthens inference.

What evidence should I request from a provider?

Questions, methods, dates, denominators, captures, interventions and comparison. Ask what the findings still do not establish.

How murmur works

murmur works through a SWAS model, combining proprietary measurement software with specialists who diagnose, plan and execute GEO to grow Share of Voice (SOV). Eligible contract models may include an SOV-growth guarantee, with the metric, baseline, period, scope and conditions agreed in writing. This does not guarantee an individual answer, a fixed position or sales.

Author: Mateus Gomes, murmur operator. This article explains our approach and discloses our commercial interest.

See also