13% of Audited Sites Block an AI Crawler — Correcting Our Own 40.8%
We re-read every robots.txt in our sample and found 2 defects in our own parser. The corrected rate is 6 of 45 audited sites — 13.3%, not the 40.8% we first published. The corrected count combines training and search crawlers; it is not an AI-search exclusion rate.
Updated August 31, 2026. The version of this article published on August 20 led with "20 of the 49 sites we audited — 40.8%." That number was wrong twice over. 40.8% was the share of individual audit runs that failed the check, not the share of distinct sites, and multiplying it back onto a site count produced a "20 sites" that nobody had ever counted. Worse, when we went back and re-read every robots.txt in the sample against the specification, 2 defects in our own parser turned out to be manufacturing half of the failures. The corrected figures are below, and the correction is now the more useful half of this article.
On August 31, 2026 we re-fetched the file from all 48 distinct domains audited in our current engine window, and 45 of them served one. Of those 45, 6 sites — 13.3% — shut at least 1 major AI crawler out of the entire site. At the opposite extreme, 8 sites have no rules that apply to AI crawlers at all — not even a wildcard directive for them to obey. That is 17.8% of the sample.
That leaves 31 sites — 68.9% of the sample — with a robots.txt that states a policy an AI crawler will actually read and blocks none of them sitewide. The count does not reveal how many owners deliberately reviewed these policies, and a restriction on a training crawler alone does not establish loss of search visibility.
| Blocks at least 1 AI crawler sitewide | 6 | 13.3% |
| No rules applying to AI crawlers at all | 8 | 17.8% |
| Explicit rules, no sitewide blocks | 31 | 68.9% |
Methodology. The population is every distinct domain with a completed audit in our current engine window, which runs from May 23 to August 31, 2026. That is 48 domains across 228 report runs. The stored verdicts carry the defective parser's answer, so the table above is not the stored tally. It is a same-day re-fetch of those domains' live robots.txt files, read on August 31 and evaluated with the corrected parser. 3 of the 48 no longer serve a robots.txt at all, which is why the denominator is 45. The sample is self-selected — these are sites somebody chose to run an audit on — and it skews small-to-mid-size. Treat the rates as directional for the long tail of the web rather than the web at large. Domains are not named.
The correction is part of the evidence
Re-reading the files exposed incorrect handling of empty disallow values and the relationship between named crawler groups and wildcard rules. RFC 9309 supplies the public interpretation: an empty path is ignored, and applicable named groups take precedence over the wildcard group. Groups for the same crawler can need to be combined; file order alone is not a reliable policy.
The defects were corrected and regression tests preserve the cases that exposed them. In this sample, 5 sites previously labeled blocked were not, while 1 previously labeled clear was blocked. The corrected result is a dated observation of fetched files. It does not establish today's access or a web-wide prevalence rate.
robots.txt is only one access surface. A CDN challenge, login wall or network refusal can prevent retrieval even when the file permits it. Conversely, a deliberate training restriction is not proof that a site is invisible in AI search. The 13.3% figure combines crawler purposes and cannot be presented as an AI-search exclusion rate.
Search access and training permission are different decisions
OpenAI's bot documentation distinguishes OAI-SearchBot, used for search discovery, from GPTBot, used for potential model training. Their controls are independent. ChatGPT-User represents user-directed visits and has a different role again. A policy review should name the purpose it intends to allow or restrict.
Anthropic documents separate crawlers for training, search and user retrieval too. Do not turn the broad phrase AI bot into a single permission toggle unless that is truly the owner's intent.
Google has its own control surfaces. Our generative search guide covers its Search Console inclusion setting alongside crawl, index and snippet eligibility. Google-Extended is a separate control; its presence in robots.txt does not substitute for reviewing the Search setting.
These distinctions are commercially useful. A business can want its public documentation discovered by prospective customers while making a separate decision about training. Neither an allow rule nor an inclusion setting guarantees an answer will cite the page. They establish part of the conditions under which discovery can happen.
Review one important URL against the policy you intended
Choose a page a prospective customer should be able to find: a useful guide, product explanation or developer example. Record the provider and crawler purpose, the applicable policy, and the response actually delivered. Keep a timestamp and distinguish a real verified provider request from a test that merely uses its user-agent name.
If the owner wants search discovery and the relevant search crawler is unintentionally disallowed, repair that specific conflict. If a training block is deliberate, record it as deliberate. Do not remove it solely to make a broad bot score greener.
A missing named rule is not automatically a defect. A wildcard rule may govern the crawler; if no group applies, the robots standard does not treat that absence as a prohibition. Explicit rules can document intent, but more lines are not inherently better access policy.
After changing the file, inspect the live response. A provider-managed file or cache can differ from the repository copy. Then verify the important page itself, including the final destination after any redirects. The indexability conflicts article explains why permission to crawl and permission to index are separate observations.
Measure whether access becomes useful visibility
Preserve the policy and response evidence before looking for outcomes. Google generative-AI impressions, provider citations, referral visits and useful product actions each answer a different question. Record their source and reporting window instead of rolling them into an invented AI visibility total.
For SEOReport, the strongest public evidence is an affected page, the owner's intended behavior, a concrete repair and a repeatable verification result. Readers can evaluate those facts without needing the private detection recipe. The AI search readiness guide connects these observations into a practical review.
This correction changes the lesson from the original headline. Our sample supports a small, dated count of crawler restrictions. The useful work is deciding which restrictions match the owner's intent, checking the delivered page, and measuring the outcomes the chosen search surface actually reports.
Get the complete diagnosis of your site
An evidence-backed report and a prioritized action plan, on a plan with monthly credits.