How Perplexity Selects and Ranks Its Sources
Perplexity retrieves through its own crawler, PerplexityBot, alongside third-party search, then answers with inline numbered citations. It cites more sources per answer than most assistants, which makes it the surface a new domain enters first. Selection favours specific checkable claims over general authority. This page covers the crawler requirements and what makes a claim attributable.
Four questions decide a Perplexity citation. Each one is answered below, in the order retrieval resolves them.
- Retrieval path — the two source streams feeding one answer
- Citation density — the measured count of sources per answer
- Crawler access — the three conditions a page meets before PerplexityBot reads it
- Claim shape — the three properties that make a fact citable
Questions
Common questions
Every answer below ships in the raw HTML, so an assistant reading this page without running a script still receives it.
Does PerplexityBot collect training data?
No. Perplexity's documentation states that PerplexityBot is not used to crawl content for AI foundation models, and that Perplexity-User collects nothing for foundation-model training either.
Does blocking Perplexity-User in robots.txt stop the fetch?
No. Perplexity-User ignores robots.txt rules, so a disallow line has no effect. Control over that token runs at the firewall, against the IP ranges published at perplexity.com/perplexity-user.json.
Which assistant does a new domain enter first?
Perplexity, on citation count alone. Every other surface ranks on a different axis, and the least controllable is set out in How Meta AI Surfaces Sources. ChatGPT resolves sources through a separate index — more on why bing decides your chatgpt visibility.
How does a business get cited by Perplexity?
Three steps run in order: open the crawler path, move the claim into server-rendered HTML, and attach a figure to its basis. Those steps form the spine of our optimization work.
Sources
- Perplexity, Perplexity Crawlers —
docs.perplexity.ai/guides/bots, checked 21 September 2026. Documents two user agents. PerplexityBot: stated purpose is to surface and link websites in search results on Perplexity, with the explicit statement that it is not used to crawl content for AI foundation models, and a recommendation to allow it in robots.txt. Perplexity-User: fetches a page in response to a user question, collects nothing for foundation-model training, and ignores robots.txt rules because a user requested the fetch. Published stringsPerplexityBot/1.0andPerplexity-User/1.0; IP ranges atperplexity.com/perplexitybot.jsonandperplexity.com/perplexity-user.json. - Qwairy, Perplexity vs ChatGPT: AI Citation Study (Q3 2025) —
qwairy.co/blog/provider-citation-behavior-q3-2025, checked 21 September 2026. 118,101 AI-generated answers, 669,065 citations, eight providers, five reported with per-answer averages: Perplexity 21.87, Google AI Overviews 17.93, Gemini 17.11, ChatGPT 7.92, Microsoft Copilot 2.47. - Vercel and MERJ, The rise of the AI crawler, 17 December 2024 —
vercel.com/blog/the-rise-of-the-ai-crawler. 24.4 million PerplexityBot fetches in one month; PerplexityBot named in the set of crawlers recorded executing no JavaScript. - Jon Bratseth, chief executive of Vespa.ai, Perplexity builds AI Search at scale on Vespa.ai, 15 April 2025 —
blog.vespa.ai/perplexity-builds-ai-search-at-scale-on-vespa-ai/. Perplexity's in-house retrieval stack on Vespa, indexing billions of web pages. - Mark Sullivan, Perplexity AI CEO Aravind Srinivas on plagiarism accusations, Fast Company, 21 June 2024 —
fastcompany.com/91144894/perplexity-ai-ceo-aravind-srinivas-on-plagiarism-accusations. Srinivas: "We don't just rely on our own web crawlers, we rely on third-party web crawlers as well."
Next step
Find out which assistants name you today.
The audit runs 40 checks across five categories and queries seven assistants for your citation baseline. Findings in ten business days.