CCBot: Common Crawl's crawler, user agent and robots.txt rule
CCBot builds Common Crawl's open web archive. How it reads robots.txt and Crawl-delay, its Opt-Out Ledger, and how to verify it by reverse DNS.
- Operator
- Common Crawl
- robots.txt token
CCBot- User agent
CCBot/2.0 (https://commoncrawl.org/faq/)- Purpose
- Open web archive
- Honours robots.txt
- Yes, its operator states it does
- Published IP ranges
- https://index.commoncrawl.org/ccbot.json
- Reverse DNS
*.crawl.commoncrawl.org- Official documentation
- Common Crawl: CCBot
- Verified on
- 2026-10-05
What CCBot does
CCBot is the crawler of Common Crawl, which says it provides a copy of the internet to researchers, companies and individuals at no cost, for the purpose of research and analysis. [2]
Common Crawl says that, currently, CCBot does not execute JavaScript or use cookies, so it sees the HTML a server returns before any script runs. [2]
Allow or block CCBot in robots.txt
Add one of these groups to the robots.txt at the root of each host (each scheme and port counts separately) it should apply to. A crawler that finds one or more groups naming its token follows those groups, merged into one, and ignores the User-agent: * group, so repeat in them any * rules you still want applied (RFC 9309, section 2.2.1).
User-agent: CCBot
Allow: /User-agent: CCBot
Disallow: /What blocking CCBot changes
Common Crawl says CCBot checks robots.txt first and fetches a page only if crawling it is allowed, obeys Crawl-delay, and will periodically check whether robots.txt has been updated, so a new rule applies once it is re-read. [1][2]
Separately, Common Crawl publishes an Opt-Out Ledger listing every legal opt-out request it has received, to alert users of its data to material that owners asked to have excluded. [3]
How to verify a request really is CCBot
Common Crawl warns that some crawlers falsely identify themselves as CCBot. The real one crawls from dedicated IP ranges published as a JSON list, and its IPv4 addresses resolve to hostnames under crawl.commoncrawl.org, so a forward-confirmed reverse DNS lookup alongside the list confirms it. Reverse DNS is not yet supported for its IPv6 addresses. [1][4]
A crawler policy is one part of being readable by assistants. The AI search readiness guide covers the rest, and robots.txt and crawl access covers the firewall and bot-protection rules that block crawlers before robots.txt is ever read. A full SEO report covers AI search readiness alongside the rest of a site's technical health.
Sources
- Common Crawl: CCBot, verified on 2026-10-05
- Common Crawl: FAQ, verified on 2026-10-05
- Common Crawl: Opt-Out Ledger, verified on 2026-10-05
- Common Crawl: CCBot IP ranges (JSON), verified on 2026-10-05
Get the complete diagnosis of your site
An evidence-backed report and a prioritized action plan, on a plan with monthly credits.