Check Whether Bot Protection Blocks the AI Search Traffic You Want
Verify the intended provider, the live response and the policy that controls it. Separate search retrieval, model training and user-directed access before changing bot rules.
A site owner can permit search discovery and still receive a report that a request was blocked. Resolving that apparent contradiction starts with the exact request: which provider was trying to reach which page, for what purpose, and what response did it receive?
An SEOReport analysis can supply evidence of an access problem to investigate. Treat that observation as the start of a diagnosis. The policy for a particular provider and the behavior of a particular request both matter.
Decide which purpose you intend to allow
AI services use different agents for different tasks. OpenAI documents OAI-SearchBot for search, GPTBot for potential model-training collection and ChatGPT-User for user-directed requests. Its crawler documentation explains that search and training settings are independent.
Anthropic's crawler documentation likewise distinguishes Claude-SearchBot, ClaudeBot and Claude-User. Record the intended choice for the provider and purpose you care about. Declining training is not a reason to label the site's search policy defective.
For Google, verify the Search generative AI control on the correct Search Console property. It is an account-level participation choice in addition to the site's technical state. Check inherited settings as well as explicit choices.
Compare the policy with the delivered response
Review the public robots.txt, the hosting or CDN rules and the actual page response. They answer different questions. Robots.txt declares crawler policy; an access rule can prevent a permitted request from receiving the page.
Inspect a representative important URL and record the observed final destination, status and content. A challenge page served with a successful status is still a challenge page. Conversely, a CAPTCHA widget on a working contact form does not establish that the whole page was blocked.
Our crawl-access guide helps organize the review of intended policy and affected URLs.
| An identified search provider receives a challenge | That observed request could not retrieve the intended page | Which rule caused it and whether the behavior repeats |
| A local test request receives a 403 | The test request was refused | Whether a verified provider receives the same response |
| A page contains a CAPTCHA widget | The page includes a human-verification control | Whether the requested content was actually withheld |
| Robots.txt permits the intended crawler | The declared crawl policy allows access | Whether hosting rules deliver the page |
These are illustrative interpretations. Preserve the real observation beside the actual rule investigation.
Verify provider identity before changing an exception
A user-agent label is easy to copy. Use the provider's documented identification guidance and your hosting provider's supported verification tools when deciding whether to admit traffic. A test with a copied label can reveal how that request is handled, but it cannot prove that the request came from the provider.
Prefer the smallest change that implements the intended policy. A legitimate search-access problem does not require removing every protection from a site. Keep authentication and customer-data boundaries intact, then verify that the intended public content is delivered to the intended reader.
Recheck after a hosting or policy change
Save the original observation, record the changed setting and repeat the relevant request. Include the page body in the check so that a changed status code cannot hide a challenge or empty response.
Provider defaults and account options can change. Read the current setting on the actual site rather than inferring it from a provider's announcement or a configuration file in the repository. On a managed platform, also confirm which layer serves the public robots.txt.
The wider agent acceptance test extends this review from reading a page to completing a permitted task. Search access, usable content and a successful action each need their own observation.
Observe discovery after confirming access
Once the intended reader can retrieve the page, allow for the provider's processing and reporting delays. Review dedicated AI impressions or citations where available, identifiable referral visits and useful product actions separately.
Restored access establishes a technical repair. A later citation or customer visit establishes a different outcome. Keeping both records gives the team a defensible account of what improved and where further work may be valuable.
See How Your Site Ranks
Get a free AI-powered SEO report with actionable findings and priority fixes for your website.
No signup required.