Guides

How to check whether AI crawlers can read your website

Check robots.txt, HTTP responses, raw HTML and verified bot logs, then give your client a reproducible access finding and a specific recheck.

A page can look correct in your browser while returning a challenge or an empty HTML shell to a crawler. For an agency technical lead, the useful result is a report showing which public page failed, what came back, and what needs checking next. Start with a page the client actually wants an answer engine to use.

We recommend checking permission, delivery and page content separately, then looking for verified crawler requests in the logs. A successful request from your laptop establishes what your laptop received. It does not prove that an answer engine has indexed or cited the page.

1. Choose the public page and crawler purpose

Confirm which public URLs the client wants crawled before changing any bot controls. Use a product, pricing or documentation page that answers an important buyer question, and get authorization to test the site. Keep authenticated pages and confidential material outside this check.

An AI crawler fetches pages for a particular purpose, which may be search or model training. OpenAI documents OAI-SearchBot for ChatGPT search and GPTBot for training, with separate controls for each. Use its current crawler documentation to choose the relevant bot and verify its published user-agent and IP information. Do not change a training policy just to diagnose search access.

Record the URL, intended crawler, test time and expected page content. Repeat the procedure for another page when it uses a different host, application or access policy. A homepage result does not describe every URL on the site.

2. Read the deployed robots.txt

Fetch the live file from the same host as the page, rather than relying on a CMS setting or repository file. The host's deployed response is what a crawler encounters. The examples below use the reserved example.test domain, so replace it with an authorized public host before running them.

curl --silent --show-error --location --max-time 20 \
  --dump-header robots.headers --output robots.txt \
  'https://example.test/robots.txt'

Inspect the status and redirect headers, then read the file. If the body is a login page or a challenge, record that response before trying to interpret it as a robots policy. For a valid file, find the rules applying to the intended crawler and the exact path you are testing.

For example, this fictional policy allows one public docs path for search while restricting training:

User-agent: OAI-SearchBot
Allow: /docs/
Disallow: /

User-agent: GPTBot
Disallow: /

The more specific path rule allows /docs/ for OAI-SearchBot in this example. Check the whole file for other matching groups and rules, especially if a CDN adds managed content. RFC 9309 defines group matching, merging of matching groups and the longest matching path rule.

If the response also contains a Content-Signal declaration, compare its uses with the client's intended policy using our deployed Content Signals guide. Keep that finding separate from whether the HTTP request succeeded.

3. Compare ordinary and simulated crawler requests

Save the response body as well as the headers, because a 200 response can still contain a challenge or login screen. Use a GET request so you inspect the page content rather than only a HEAD response. Run these requests without login cookies or authorization headers.

PAGE_URL='https://example.test/docs/getting-started'

curl --silent --show-error --location --max-time 20 \
  --dump-header ordinary.headers --output ordinary.html \
  --write-out 'HTTP %{http_code}\nFinal URL %{url_effective}\n' \
  "$PAGE_URL"

curl --silent --show-error --location --max-time 20 \
  --user-agent 'OAI-SearchBot' \
  --dump-header simulated.headers --output simulated.html \
  --write-out 'HTTP %{http_code}\nFinal URL %{url_effective}\n' \
  "$PAGE_URL"

The second request uses the bot's name as a diagnostic token. To investigate a rule matching the full user-agent string, repeat it with the current string from the operator's documentation and record exactly what you used. Neither request comes from OpenAI's crawler infrastructure.

Compare the status, redirect chain and body:

Observation Next check
Ordinary request succeeds, simulated request returns 403 Inspect the security event and matching CDN or web application firewall rule
Both requests reach a login page Confirm this is an intended public URL and investigate the authentication redirect
200 response contains a challenge Inspect bot controls and the challenge response instead of marking the page readable
Request times out or has a network error Record an unresolved test and investigate connectivity before reporting a block
Both return the expected page Continue to the HTML and verified-log checks

Cloudflare's Block AI Bots documentation describes edge controls that can affect access independently of the site's robots file. Ask the security owner to inspect the rule responsible for the observed response. A simulated request can be treated differently because of its IP address, headers or client behavior, so investigate the event before changing a rule.

4. Check what is present in raw HTML

Read the saved HTML for the actual information the page should supply. Find the heading and a distinctive passage, such as a product condition or a setup prerequisite. Check both saved bodies, not only the browser's rendered page.

If the browser shows that passage but the response contains only a JavaScript application shell, document which content is absent. Consider serving that important content in the initial HTML, then repeat the request. This test does not establish the rendering capabilities of every AI client, but it shows what a client receives without executing JavaScript.

Maverank's readiness audit checks crawler policy, simulated fetchability and raw-HTML content as separate findings. It also distinguishes checks that could not be measured. Keep that distinction in the client report: a timeout is an investigation, not evidence that the client deliberately blocked a crawler.

5. Verify real crawler requests in the logs

Use the crawler operator's current identity-verification guidance before treating a user-agent entry as a genuine bot visit. Anyone can send a bot name in a header, as the simulation above demonstrates. For OpenAI, check the source IP against the published ranges for the specific crawler.

Review CDN or edge logs as well as origin logs when a security layer sits in front of the site. A blocked edge request may never reach the origin. Use the trusted client-IP field from your logging stack, rather than an arbitrary forwarded header supplied by the requester.

For a verified request, record the time, requested path, status and any blocking or challenge event. If there is no matching request in the available logs, report that you have not observed a verified visit in that window. Log retention, sampling and crawl timing can all limit that observation.

6. Give the client an evidence table and a recheck

Report each layer separately so the person fixing the problem knows where to start. Here is a fictional finding for an authorized public docs page:

Layer Observed evidence Action
Policy /docs/ allowed for the intended search bot in deployed robots.txt Preserve the approved policy
Simulated delivery Ordinary request 200, simulated request 403 Security owner checks the matching edge event
Page content Expected prerequisite present in ordinary raw HTML Recheck the same passage after the access change
Real bot access No verified visit found in the available log window Check later visits and retain the window limitation

Change only the control that conflicts with the client's intended access policy. Do not disable the firewall or open private routes to make a test pass. After the change, fetch the deployed robots file and the same page again, then inspect subsequent verified requests when they are available.

Use the finding to turn an AI readiness audit into a client brief, with an owner and the exact recheck. If you use Maverank for agencies, start with one authorized public site and investigate the relevant readiness finding before expanding the work.

Frequently asked questions

Does a successful simulated crawler request prove ChatGPT can cite the page?

No. It shows that your test received the page from that network with those request headers. A verified bot visit provides evidence of actual access, while indexing and citation remain separate observations.

Should we allow every AI bot to fix access?

No. Agree the public pages and crawler purposes first, then investigate the control affecting the intended crawler. Keep authentication and protections for private content in place.