GEO can support a clinic’s discovery when patient questions are local and specific; the decision to invest should start with a question set, a baseline and an SOV target defined for that operation. Results from companies that market to clinics are not automatically results from clinics. This guide therefore distinguishes internal examples, client evidence shared with permission, and hypotheses that still need to be measured in your market.

I am Mateus Gomes, operator of murmur.marketing. Our operation combines proprietary software for measuring SOV with specialists who define and execute GEO strategy. The internal measurement from 2026-08-06 is a historical baseline from before the current guide library; it does not assess today’s SWAS strategy and should not be extrapolated to clinics. Where an eligible contract includes an SOV-growth guarantee, scope, baseline, timeframe and conditions are defined in the contract.

This page connects to What is generative search and how does it change marketing in 2026?. Its scope is health, and the boundary below bears repeating.

Summary

  • I have zero measurements of a clinic, practice or insurer. Anyone presenting a Portuguese health-vertical study should show the denominator and execution date before the chart.
  • I do have Fly Vet and Fly Med, the operator’s own companies: agencies serving veterinary clinics and medical specialists.
  • At Fly Med, ChatGPT first cited a directory page at T+5 days: two pages in one question, with `?utm_source=chatgpt.com` in the URL.
  • On the same day and with the same method, sister company Fly Vet had zero cited directory URLs.
  • Where a veterinarian describes the problem in first-person language, Fly Vet appears 2 times in 80 in ChatGPT and 0 in 80 in Claude.
  • In the June 2026 Fly Vet sample, engines differed: ChatGPT recommended it in 30 of 100, Claude in 41 of 100, while Google AI Overview did not recommend it in that sample and used its site as a source in 81 of 100.
  • For clinics, measure whether content is retrieved, cited, mentioned or recommended; these agency examples do not substitute for clinic-specific measurement.

The boundary, stated plainly

Fly Vet and Fly Med are my companies. They are marketing agencies: the former serves veterinary clinics, the latter medical specialists. Their numbers describe two B2B agencies in questions clinic owners ask before hiring marketing. They do not describe a clinic in questions patients ask before booking an appointment.

The markets, buyers and generative-engine behavior differ; it varies even between two question families in the same market. What transfers is the mechanism and the measurement method. What does not transfer is the number. I cannot name a client without recorded consent, so the named material here is limited to the operator’s agencies, with dates, denominators and scope stated.

The Fly Med case: citation at T+5 days

This is the favorable result, and it is intentionally small. Fly Med’s content directory was published on 2026-05-22. In a probe run on 2026-05-27, five days later, ChatGPT cited two pages from that directory in a single question, first and third in its references, with `?utm_source=chatgpt.com` in the URL. The repository has the screenshot and capture date.

Five days is quick enough to surprise anyone coming from traditional search. The next number matters more.

The Fly Vet case: zero, on the same day, with the same recipe

In the same probe on the same date, the sister company had zero cited directory URLs. Same method, team, content structure and market type. Zero. Publishing only T+5 would create a time expectation my own material cannot support. In a third, unnamed client subdomain, first citation came at T+28 in Perplexity and T+30 in ChatGPT. T+5, T+28, and zero: anyone promising a timeframe is guessing.

The full probe needs an honest note: 32 mentions in 59 successful captures (of 65 attempted), a number someone could sell as “49% mention.” Only 2 of those mentions were directory URLs. Brand mention is not content working.

The pain-language blackout: where a clinic loses most

The Fly Vet question universe contains a family in which the buyer describes the problem in first person: for example, “my clinic does not show up on Google.” The denominator is 80 records per engine (40 questions × 2 repetitions), campaign completed by 2026-06-30, using the `citation_kind` rule. The result: ChatGPT 2 of 80; Claude 0 of 80.

The same company wins comparative and category questions in other families. It wins where the buyer already knows what they want and disappears where the buyer describes the problem in their own words.

For a clinic, this asymmetry is the central lesson. Patients rarely type a procedure name with medical-record precision. They type a symptom, fear, schedule constraint and neighborhood. Content covering only technical vocabulary covers the question family with the greatest contest and the buyer furthest along, while leaving empty the family where a decision begins.

Four engines disagree about the same clinic — and the same agency

For Fly Vet, across 100 questions in June 2026, one repetition per question under `citation_kind`: ChatGPT recommended it in 30 of 100, Claude in 41 of 100, while Google AI Overview did not recommend the agency in that sample and used its site as a source in 81 of 100.

Those are incompatible behaviors about the same company in the same month. Measuring only Google yields an apparently excellent 81 of 100, but it describes bibliography, not an answer. Measuring only Claude yields a different conclusion. One engine is one engine.

Scope limit: the engine called `google` here is AI Overview, the generated box at the top of search. The Gemini app is not measured in this pipeline; no collector exists for it.

“Done but not winning”: the pattern a clinic must recognize

Fly Vet has a published content cluster: 324 source articles, all live, and a homepage claiming a leading category position. Yet in its June 2026 campaign, ChatGPT cited EvolueVet in 28 of 100 answers; of those 28, Fly Vet was recommended in 10. In 8 questions asking for the “best marketing or traffic agency for veterinary medicine,” Fly Vet was absent and EvolueVet appeared first. These figures are recounts by question identifier in the verdict file, not a report-line count.

Content is done; the game is lost. When that happens, the bottleneck has stopped being content and has become entity: the engine does not resolve the brand as a distinct thing and fills the space with a brand it does resolve. I have the extreme version myself. On 2026-08-06, in ten questions that already contained my name × 4 engines, denominator 40, engines cited a domain with “murmur” in its name in 33 captures, but only 5 were mine; in the other 360 captures, no “murmur” domain appeared.

Name cited in “veterinary marketing agency in Brazil”What the measurement recordsDenominator and ruleWhat the number is not
EvolueVetMost cited in the category: 28 answers100 ChatGPT questions, June 2026; recount by question identifierIt is not a performance measurement of that company; it counts a name in an answer
Fly Vet (the operator’s company)Recommended in 10 of the 28 answers where EvolueVet appears; absent in the 8 “best agency” questionsSameFinding from this sample, not a direct score between the two
Agência Pet · VetConcept · MarketVetNamed in the category, without a published countSameNo count is not no presence
Miumarqué · Gvet · VetDesenvolveNamed in the category, without a published countSameSame
`murmur.marketing`Absent in the historical baseline, before the current guide library800 captures, 2026-08-06Not a current SOV measure or a clinic result

How to test this in your clinic in one afternoon, without hiring anyone

  1. Write 10 to 25 questions as your patient would write them: symptom, fear, schedule constraint and neighborhood, not medical-record vocabulary. Include first-person pain questions; a historical Fly Vet campaign recorded 2 of 80 in ChatGPT and 0 of 80 in Claude for that family, but those figures are not a current clinic benchmark.
  2. Run each question five times in each of four engines, in a clean conversation without history. Five times because language models are stochastic; one run is invalid, with observed variation of 10% to 34% between identical executions.
  3. Screenshot every capture. What appeared on screen is the only evidence that survives discussion. I have 800 of 800 screenshots; one failed and stayed in the denominator because dropping a capture invites cherry-picking.
  4. Classify four labels, not one: recommended, source, mentioned, absent. Collapsing them into “mentions” almost made me publish “49% mention” when only two mentions were content working.
  5. Measure noise before effect. In 387 questions rerun three days apart with no intervention, 98% of verdicts repeated. Movement above 3 to 4 percentage points is real; below that is variance.

If most questions return without naming any clinic, you are looking at an empty chair. Claiming an empty chair is the only situation where this investment is genuinely inexpensive.

What I do not know about health, and will not pretend to know

I have no denominator by health vertical. The 800 captures from 2026-08-06 concern my own category, not clinics. No number on this page describes what AI tells a patient.

The 2026-08-06 baseline is not a treatment study. It measured `murmur.marketing` before the current guide library and cannot assess today’s SWAS strategy or establish what works for clinics. For a clinic, start with its own baseline, question set and SOV target; any guarantee applies only to eligible contract models under agreed terms.

External context is required. Olivier Martinez’s arXiv 2607.14035, submitted 2026-07-15, reviews 45 studies from November 2023 through July 2026 and concludes that none of the reviewed techniques demonstrates a stable, longitudinal, cross-platform causal effect; it includes slices where GEO-style rewrites reduce retrieval. The founding literature also exists: Aggarwal et al., KDD 2024, arXiv 2311.09735, and its 10,000-query GEO-bench.

There is a technical condition that defeats more practices than expected. Vercel reported in The rise of the AI crawler that none of the major AI crawlers renders JavaScript; they download scripts without executing them: 11.50% of OpenAI-crawler requests and 23.84% of Anthropic-crawler requests. A modern-builder clinic site whose content arrives by script is an empty page to GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot.

Frequently asked questions

Do you have a clinic or private-practice case to show?

No. I would rather state that than adapt a case from another market. I have named results from my own Fly Vet and Fly Med agencies, serving veterinary clinics and medical specialists respectively. They are suppliers to clinics, not clinics. The mechanism and measurement method transfer; the numbers do not.

My clinic site already appears on Google. Is that enough for AI?

Not necessarily. Generative engines use search indexes to retrieve documents, but AI crawlers do not execute JavaScript: they read the HTML delivered by the server and leave. Many clinic sites load their content after a script runs, so the page exists for a visitor and does not exist for a crawler. That can be checked in minutes, before spending on copy.

My site has a page for each procedure. Is that already one page per question?

Almost never. A procedure page is organized by the service name, the medical-record vocabulary; a question page is organized by the wording a patient arrives with. Read every existing title and ask whether someone would type it. If it is a procedure name, it answers a catalog query while the generative engine is retrieving a document for something else. Renaming a title without changing the scope does not solve it: the change is one page per wording, not one page per service.

What is the most common content error in health?

Covering only technical vocabulary. In the Fly Vet campaign completed on 2026-06-30, first-person pain questions returned 2 of 80 in ChatGPT and 0 of 80 in Claude. This is an agency example, not a clinic benchmark. Patients type symptoms, fears, schedule constraints and neighborhoods, not a procedure name.

My clinic serves one neighborhood. Does the afternoon test change?

The design does not change: 10 to 25 questions, five repetitions, four engines, a screenshot per capture and four labels. The question wording must add neighborhood and schedule constraint to the symptom, because that is how patients write. I have no measurement of local clinic questions and will not extrapolate mine. I can only point to where this family disappeared at Fly Vet: the buyer’s own-problem wording returned 2 of 80 in ChatGPT and 0 of 80 in Claude, in a campaign completed by 2026-06-30.

Who wrote this, and disclosure of interest

Mateus Gomes, operator of `murmur.marketing`, a GEO-only operation in Brazil. Fly Vet and Fly Med are my companies, so I can publish their figures, including their bad ones. No client is named without recorded consent.

For context on the historical evidence and my interest, the 2026-08-06 baseline recorded zero in 800 captures across four engines. The text rule’s 88 hits in 800 all belong to the ten questions that already contained the word; outside them, 0 of 712. This predates the current guide library and is not a current visibility score.

Conclusion

GEO is worth evaluating for a clinic when measurement reflects what patients actually ask — symptom, fear, constraint and neighborhood — and the site is readable by crawlers. The examples here are clearly bounded: Fly Vet and Fly Med are marketing agencies, not clinics, and their results cannot be projected onto a clinic. They illustrate why measurement must separate recommendation, citation, mention and absence. Start with the clinic’s own baseline and track SOV as strategy and execution progress; guarantees are available only in eligible contract models with scope, baseline, timeframe and conditions defined in the agreement. The one-afternoon test above helps establish whether there is a measurable opportunity.

See also