Appearing in Perplexity starts with being retrievable: it shows the sources it used, so the path from your page to the answer is visible — and that makes it the easiest of the four engines I measure to diagnose. That does not make it easier to win. Visible citations can mislead: your brand may be in the source list without being named in the answer itself.
I write from murmur.marketing, a GEO (Generative Engine Optimization) practice in Brazil,
operated by Mateus Gomes. On 2026-08-06, I ran 100 questions in Portuguese against four engines,
including Perplexity, producing 800 browser captures with screenshots for all 800.
Summary
- Perplexity exposes its sources, making diagnosis a matter of direct inspection. With the other engines you infer; here you can see.
- This creates the opposite trap: being a source is not being recommended, and the visible list makes the first state look like the second.
- Its crawler is PerplexityBot. In an nginx log covering 3.6 days, it made 62 requests, compared with 327 from Googlebot and approximately 184 from OpenAI’s three user-agents combined.
- ⚠️ Counting crawlers by user-agent is misleading. In a different, 24-hour log, the raw count showed 6 Perplexity hits. Checking the real IP addresses left zero: all six were spoofed.
- The engine answers without a login, making it the cheapest to measure honestly — and also the easiest to measure badly with a single question.
- For the question that started this project, repeated five times in each engine, the most frequently cited name appeared in 4 of 5 Perplexity runs and 0 of 5 Claude runs, on the same day.
What differs from ChatGPT
The four engines I measure — ChatGPT, Claude, Perplexity and Google AI Overview — do not play the same game. Treating them as one is the most common beginner’s mistake. The relevant difference here is where they get their material.
| Engine | Where its material comes from | What you can observe |
|---|---|---|
| Perplexity | Its own search, with sources displayed in the answer | The list of URLs used: direct diagnosis |
| ChatGPT with search | Grounding in Bing’s index | Citations with ?utm_source=chatgpt.com in the URL |
| Claude | Its own search, with more sparing citations | Little: it cites least often, and identifying that requires measurement |
| Google AI Overview | Google’s own index | The AI block above organic results, with links alongside it |
The practical consequence is where to intervene. For ChatGPT, Bing’s index is the bottleneck, making Bing Webmaster Tools an opportunity almost nobody in this market uses. For Perplexity, the bottleneck is whether your page can be retrieved and answers the question in a passage that survives extraction. There is no third-party index to court.
Four actions that hold up
None of these is specific to Perplexity, and that is the honest observation: there is no Perplexity technique. It is the same sequence as elsewhere, with more visible feedback.
1. Be accessible, and prove it in the log
The robots.txt at the domain root needs to allow PerplexityBot and declare the sitemap.
An easily missed detail: crawlers read the root robots.txt, and only that file. Nobody fetches
one published in a subdirectory. Also, AI crawlers do not execute JavaScript: if the served
HTML is empty and the content appears only after hydration, there is nothing to retrieve.
There is no second chance.
The way to establish whether this worked is your server log, not a vendor’s tool. On a client’s subdomain, over 3.6 days and 1,074 requests, PerplexityBot made 62 requests, Googlebot 327, and OpenAI around 184 across OAI-SearchBot, GPTBot and ChatGPT-User. Bingbot made exactly 1. AI crawler traffic is already in the same order of magnitude as Googlebot’s, and the distribution is not what intuition would suggest.
2. Answer at the start instead of building up to it
Perplexity assembles answers from passages. A warm-up paragraph before the answer is material that will not survive extraction. Put a complete, factual statement in the first line after the heading. The same principle is explained in How to structure content to be cited by AI.
3. Name the entity alongside the fact
The most common failure is not a failure to retrieve: it is being retrieved and paraphrased without credit. The engine describes your content without saying who produced it. The remedy is in the writing: put the company name in the same sentence as the finding, not in a signature at the bottom of the page.
4. Give numbers with denominators
The three effects Aggarwal et al. measured against GEO-bench — 10,000 queries × 10 engines — point in this direction: source citations +115%, statistics +41%, and direct quotations +30%, against the benchmark described in the paper (arXiv 2311.09735, KDD 2024). This is the category’s founding literature. It is also worth recording that the critical survey arXiv 2607.14035 rates the paper level C in its own evidence hierarchy.
The trap in an engine that shows its sources
This only becomes apparent after measurement, and it is why this article deserves its own page.
Because Perplexity displays sources, it is tempting to count appearances in that list and call them results. But being in the source list and being recommended in the body are different states, requiring different remedies: a source with no body mention is a writing problem; a mention without being chosen is a positioning problem. The four states I use to distinguish these outcomes are in Cited as a source or recommended by AI.
There is also an instrumentation trap that I fell into and am correcting publicly: citation-chip text can be captured together with the answer body, depending on how the page is read. When that happens, a name that appeared only in the sidebar is counted as though it had been said in the answer, inflating the count. If your vendor measures by screen scraping, this is a legitimate question to ask.
Why crawler counts based on user-agent are misleading
This deserves its own section: it is the most costly lesson in my data about this engine, and it is a negative finding.
In a 24-hour window containing 2,256 access-log lines from a client’s subdomain, the raw
user-agent count was ClaudeBot 88 · OpenAI 10 · Perplexity 6. Cross-checking each line against
the real IP recorded in X-Forwarded-For left ClaudeBot 79 · OpenAI 3 · Perplexity 0.
All six apparent PerplexityBot hits were fake. One IP sent forged user-agents for PerplexityBot,
Googlebot, bingbot, YandexBot, DeepSeekBot and CCBot in the same second, requesting /.env,
/.git/config and /service-account.json. This was not a crawler: it was a credential scan
disguised as one.
If you count AI crawlers by user-agent, you are counting attackers. That does not invalidate the 3.6-day observation above: they are different windows and sites. Our reporting rule is to state which of the two checks was performed. That is the difference between a defensible number and a number that simply gets repeated.
The competitive context, with the counting rule stated
Context for the historical numbers: Murmur’s model combines proprietary measurement software
with specialist strategy and execution to grow SOV. A before the current guide content was published baseline on 2026-08-06 recorded
citation_kind = ausente (the original label for “absent”) in 800 of 800 captures, including all 200
Perplexity captures. This is a dated baseline, not a current performance claim. Excluding the ten questions that already named the brand in the prompt,
the result was 0 of 712.
The question that started this project — “qual a melhor empresa pra fazer GEO no brasil” (English translation: “what is the best company for GEO in Brazil”) — was run five times in each of the four engines on 2026-08-06. Counting captures in which Conversion was named in the body, using a deterministic rule against a declared list of 29 names: ChatGPT 5/5, Perplexity 4/5, Claude 0/5, Google AI Overview 0/5. The measured prompts were in Portuguese.
Consider what that means for someone seeking to appear in Perplexity: the brand is almost always there in one engine and never appears in two others, across the same five repetitions on the same day. There is no single state of “being visible in AI.” There is visibility for a (question, engine) pair, with a denominator.
Across the category, in the same 100 pairs under the same rule, 51 of 100 have no winner at all. Conversion wins 19 and is absent in 81; GeoStack wins 15 and Brasil GEO 14. The three are separated by less than 1.3 binomial standard errors, so this is not a stable ranking. The position is open, including in Perplexity.
The instability is not peculiar to my instrument. SparkToro’s research by Rand Fishkin and Patrick O’Donnell, with 600 volunteers and 2,961 runs across 12 categories, reached the same conclusion at a larger scale: brand recommendations are highly inconsistent between runs. A screenshot of one question in one engine on one day is a sample of one.
Frequently asked questions
Why is Perplexity easier to diagnose than the other engines?
Because it displays the sources used to assemble its answer, making the path from your page to the text visible without inference. In the other three engines, you work out the origin through deduction or URL traces; in Claude, which cites most sparingly, you often cannot establish it at all. The trade-off is that this visibility tempts you to count source-list appearances as recommendations, mixing two states that need different remedies.
Is there a technique specific to Perplexity that does not work for the others?
None that I have managed to isolate, and I am skeptical of anyone who claims otherwise. What changes between engines is where their material comes from — Bing’s index is the bottleneck for ChatGPT, whereas Perplexity has its own search — rather than how the page should be written. The same four things apply to all of them: served HTML that does not depend on JavaScript, an answer at the start, the entity named alongside the fact, and numbers with denominators.
How do I confirm that PerplexityBot really visited my site?
Use the server log and cross-check user-agent against IP, never user-agent alone. The difference can reverse a conclusion: in a 24-hour log I analyzed, the raw count reported six PerplexityBot visits, while checking the real IP left zero. All six came from one address sending forged user-agents for seven different bots in the same second and requesting credential files. A tool that reports crawler visits without saying whether it verified the IP is reporting the inflated count.
Does appearing in Perplexity help me appear in the other engines?
Not directly, and the evidence I have points against the intuitive answer. For the same question on the same day, with five repetitions in each engine, the most frequently cited brand in my campaign appeared in four of five Perplexity answers and none of the five Claude answers. The engines retrieve from different indexes and write according to different criteria. What transfers is not the citation but its shared causes: being accessible, being specific, and existing as a recognizable entity outside your own site.
Do I need a login to measure Perplexity?
No, making it the cheapest of the four to measure honestly. You can run an anonymous query in a clean session without history contaminating the result. That convenience does not solve variance: one query is still a sample of one, however easy it was to obtain. Measuring cheaply and measuring too little are different things, and the second produces misleading reports.
Who wrote this, and disclosure of interest
Mateus Gomes operates murmur.marketing, a Brazilian GEO operation combining proprietary measurement software with specialist strategy and execution to grow SOV. The dated 2026-08-06 baseline was collected before the current guide collection and is included to illustrate method, not as a present-day performance score.
Conclusion
Perplexity is the engine where the work becomes visible fastest because it shows its sources. For exactly that reason, it is also the easiest to misread: being on the list is not being chosen. There is no exclusive technique; the actions are familiar, with a different origin for the retrieved material. If you verify crawler visits, cross-check user-agent and IP: in my data, that check has already turned six into zero. Do not treat a screenshot of one query as a result. With five repetitions of the same question on the same day, one engine cited a brand four times and another never did. To discuss measurement for your situation, contact Mateus Gomes on LinkedIn.