Perplexity runs its own crawler and its own index, and its retrieval is hybrid rather than purely vector-based: BM25 and n-gram matching combined with domain and page trust scores and PageRank-style authority signals. Because BM25 is lexical, the exact phrasing on your page matters more here than on most AI surfaces. And because the index is deliberately smaller and more curated than Google's, getting into the index at all is a bigger lever than ranking within it.
Key Takeaways
- Perplexity operates PerplexityBot and maintains its own index, supplemented by third-party crawlers. It is not a Google or Bing wrapper.
- The index is deliberately smaller than Google's, curated toward high-quality trusted sources. Inclusion is a filter, not a given.
- Retrieval is hybrid: BM25 plus n-gram retrieval, combined with domain/page trust scores and PageRank-style authority. Explicitly not clickstream-based.
- Because BM25 is lexical, exact query phrasing appearing on your page still carries weight, an old-school signal that most AEO advice has written off prematurely.
- The summarisation model operates under a hard constraint to use only retrieved documents, which is why hallucinated citations are comparatively rare and why retrieval, not generation, is the whole game.
- Trust-score mechanics and recrawl cadence are undisclosed. Anyone quoting specific thresholds is guessing.
The Assumption Almost Everyone Gets Wrong
Ask ten marketers how AI search retrieval works and nine will describe vector search. Embed the query, embed the documents, find the nearest neighbours in semantic space, done. It has become the default mental model, and it drives a whole genre of advice about writing "semantically rich" content and stopping worrying about exact keywords.
For Perplexity specifically, that model is incomplete in a way that changes what you should do.
Perplexity's retrieval is hybrid. BM25: a lexical ranking function that has been a workhorse of information retrieval since the 1990s, sits alongside n-gram retrieval and authority scoring. Semantic understanding happens, but it is not carrying the whole load.
The practical consequence: the literal words on your page still matter on this platform. Not in a keyword-stuffing sense. In a "does the phrase your buyer actually types appear somewhere on the page in natural language" sense.
What BM25 Actually Rewards
BM25 scores a document against a query using term frequency, inverse document frequency and document length normalisation. Stripped of the maths, it rewards three things.
Term presence
The query terms need to appear in the document. Not synonyms of them, not concepts adjacent to them: the terms. A page about "reducing customer churn" scores poorly against a query for "stop subscribers cancelling" under a purely lexical function, however well it answers the question.
Term rarity
Rare terms carry more weight than common ones. This is why specific technical vocabulary, product names and precise phrasings punch above their weight, and why generic content that could be about anything performs poorly.
Length normalisation
Long documents are not automatically favoured for containing more terms. Padding does not help. This is the mechanism that quietly punishes the 3,000-word SEO article that says 400 words of actual content.
What this means for how you write
Write the question the way people ask it, somewhere on the page, in a sentence that reads naturally. In a heading, in an opening line, in an FAQ. This is not a return to exact-match keyword density. That would fail the length normalisation and quality filters. It is the much simpler discipline of not being so abstract that the actual query terms never appear.
I have watched perfectly good pages fail to retrieve on Perplexity for queries they clearly answered, because the page used the brand's internal vocabulary throughout and never once used the customer's.
The Small Index Changes Your Entire Priority Order
This is the part that reorders strategy.
Google's index is enormous and inclusive; almost anything crawlable ends up in it. Your competition is therefore about ranking, beating the other million pages that are also indexed.
Perplexity's index is deliberately smaller, curated toward sources judged high quality and trustworthy, supplemented by third-party crawlers. Inclusion is a decision, not a default.
Inclusion is the bigger lever
If your domain is not in the index, no amount of on-page optimisation will surface you. You are not ranking 47th. You are not present. Every hour spent tuning a page that is not indexed is an hour wasted.
So the first question is not "how do I rank on Perplexity." It is "am I in Perplexity's index at all."
This is genuinely the opposite of the Google playbook, where inclusion is assumed and ranking is the whole job.
How to check
The practical test: query Perplexity for content that only exists on your site: an unusual phrase from one of your pages, a proprietary framework name, a specific data point you published. If Perplexity can surface it, you are in. If it consistently cannot find content that exists nowhere else, you have an inclusion problem, not a ranking problem.
Run this across several pages before concluding anything. A single miss could be a retrieval quirk; a consistent pattern across ten unique phrases is a signal.
If you are not in
Allow PerplexityBot in robots.txt and verify no edge or WAF rule is blocking it: the same edge-layer problem that silently blocks OpenAI's crawlers blocks this one. Then work on the signals a curated index plausibly uses: genuine external references from established sites, a clear and consistent domain identity, and content that is demonstrably not thin.
I say "plausibly" deliberately. Perplexity has not published its inclusion criteria.
Trust Scores and Authority: Documented in Kind, Not in Detail
Perplexity's retrieval incorporates domain and page trust scores and PageRank-style authority. What is documented is that these exist and are used. What is not documented is how they are computed, what thresholds apply, or how quickly they update.
What we can reason about
PageRank-style authority is a link-graph concept. Links from established, trusted domains contribute more than links from anywhere else. This is not a new idea and it has not stopped mattering.
Domain-level and page-level scores being separate implies a strong domain does not automatically carry a weak page, and a strong page on an unknown domain faces a harder path. Both halves need work.
What is explicitly not used
Clickstream data. Perplexity's retrieval is not built on user behaviour signals of the kind that inform some other systems. That is a meaningful architectural statement: you cannot influence Perplexity retrieval through engagement metrics, dwell time, or traffic-shaping. The signals are content, links and index inclusion.
What nobody knows
Recrawl cadence is undisclosed. So is the exact weighting between lexical, authority and semantic components. If you read a post claiming Perplexity recrawls every N days, or that trust scores update on some schedule, ask where that number came from. In every case I have checked, the answer is that someone made it up.
The Hard Grounding Constraint
Perplexity's summarisation model operates under a hard constraint: use only the retrieved documents. It does not free-associate from parametric memory and dress the result in citations.
Why this matters practically
It means the citation is not decorative. If your page is not retrieved, you are not in the answer. There is no path where the model "knows about you" and mentions you anyway. Retrieval is the entire funnel.
It also means the answer is only as good as what was retrieved, which is why Perplexity's output quality tracks so closely with whether the index happens to contain a good source on the topic. Gaps in the index show up as vague answers.
The implication for content design
Write pages that survive extraction. A page whose core claim is only clear after reading 800 words of build-up is a page the model has to work to summarise. A page that states the answer plainly near the top, then supports it, gives the model something clean to lift.
This is the same discipline that makes for good writing generally, which is a recurring theme in AEO work, most of the wins are quality wins with a technical label attached.
A Perplexity-Specific Checklist
Different enough from a general AEO checklist to be worth its own list.
Verify PerplexityBot access
Robots.txt, CDN rules, WAF. All three. Log-check that PerplexityBot is actually reaching you, not just theoretically permitted.
Test index inclusion with unique phrases
The unique-phrase test above, run across at least ten pages. This is your single most informative diagnostic.
Audit for query-phrasing presence
Take your twenty highest-value questions, phrased as your customers phrase them. Check whether those exact phrasings appear anywhere on the relevant pages. In my audits, roughly half do not.
Lead with the answer
Direct answer near the top of each page, in plain language, in one or two sentences that stand alone.
Build genuine external references
Authority signals in a link-graph system come from links. Digital PR, original data, practitioner commentary that other people want to cite. Slow, unglamorous, still works.
Do not chase engagement metrics for this platform
Clickstream is explicitly not part of retrieval here. Time spent optimising engagement signals for Perplexity is time misallocated.
Where I Think Most Advice Goes Wrong
Two specific errors.
First, the vector-only assumption. Advice that tells you keywords no longer matter is overcorrecting from Google's semantic improvements and generalising it to systems that work differently. On a BM25-inclusive system, phrasing matters.
Second, treating Perplexity like a smaller Google. The small curated index inverts the priority order. If you run a Google playbook here, assume inclusion, optimise for rank, you will spend your effort in the wrong place entirely.
For general search behaviour, the ongoing reporting at Search Engine Land and the technical breakdowns at Ahrefs are useful counterweights to vendor content. For Perplexity specifically, the primary sources are Perplexity's own technical statements, treat everything else as interpretation.
Frequently Asked Questions
Does Perplexity use Google's index?
No. Perplexity runs its own crawler, PerplexityBot, and its own index, supplemented by third-party crawlers. It is not a wrapper on another engine's results.
Is Perplexity's retrieval purely semantic or vector-based?
No. It is hybrid. BM25 and n-gram retrieval sit alongside authority and trust scoring. Lexical matching is a real component.
Do keywords still matter for Perplexity?
Phrasing matters more here than on most AI surfaces, because BM25 is lexical. That means using the words your audience uses, naturally, not stuffing them.
Why is Perplexity's index smaller than Google's?
It is a deliberate design choice, curated toward high-quality trusted sources. The trade-off is narrower coverage in exchange for a higher baseline of source quality.
How do I get into Perplexity's index?
Allow PerplexityBot everywhere including at the edge, and build the quality and authority signals a curated index would plausibly weigh. The specific inclusion criteria are not published.
How do I check whether I am indexed?
Query Perplexity for a distinctive phrase that appears only on your site. Repeat across ten or more pages. Consistent inability to surface unique content indicates an inclusion problem.
Does user engagement affect Perplexity rankings?
Retrieval is explicitly not clickstream-based. Engagement optimisation is not a lever on this platform.
How often does Perplexity recrawl?
Undisclosed. Any specific figure you see quoted is unsourced.
Can Perplexity cite a page it did not retrieve?
The summarisation model operates under a hard constraint to use only retrieved documents, which makes this unlikely by design. Retrieval is the gate.
Does Perplexity work differently from ChatGPT search or Copilot?
Materially, yes. Different crawler, different index, different retrieval architecture, different grounding constraints. Treating all AI search as one channel is the most common strategic error in this space.
If you want a second pair of eyes on whether your site is actually in these indexes, and what to do if it is not, the door is open. I have spent 4+ years in marketing helping edtech and startup brands grow organically, including the work behind Masai School's Instagram climbing from 26K to 117K and LinkedIn from 50K to 160K. There is more of that work, and a contact form, at younusfardeen.com. Happy to look at your setup whether or not it turns into anything.