Can AI crawlers read your site? I checked mine

I tested eight AI crawler names against this site. All eight got in. The thing that was blocked was my own share image, by a line in my own robots.txt. Here is the command, the results, what I fixed, and the five checks I now run.

Also available in 中文

A five-foot way of old shophouses in golden late-afternoon light, a row of wooden doors standing open and a single red door closed at the far end.
Every door was open to the crawlers. The one that was shut, I had shut myself.

Why I checked

Since July 2025, Cloudflare has blocked AI crawlers by default on new domains unless the owner says otherwise. This site runs on Cloudflare. I write here so that people, search engines and answer engines can find what I think, and an answer engine cannot quote a page it was never allowed to fetch.

So before doing anything clever, I wanted the dull question answered: can they get in at all?

I run this site with two Claude sessions, a Brain that plans and a Hand that builds. The audit was the Brain's job, read-only: look at the live site from the outside, change nothing, write down what is true.

The check, in one command

Every crawler announces itself with a user-agent string. You can send the same string yourself and see what your server answers:

curl -s -o /dev/null -w "%{http_code} %{size_download}\n" \
  -A "Mozilla/5.0 (compatible; GPTBot/1.1)" \
  https://yappatyih.com/writing/the-mamak-table-question/

A 200 with the full page size means that name is let in. A 403, a challenge page or a much smaller size means something in front of your site is turning it away.

Here is the run from 2 October 2026, against one post on this site:

Name Whose What it does Status
GPTBot OpenAI collects pages for training 200
OAI-SearchBot OpenAI builds ChatGPT's search index 200
ChatGPT-User OpenAI fetches a page when a user asks 200
ClaudeBot Anthropic collects pages for training 200
Claude-User Anthropic fetches a page when a user asks 200
PerplexityBot Perplexity builds Perplexity's index 200
CCBot Common Crawl open web archive many models use 200
Bytespider ByteDance collects pages for its models 200
Google-Extended Google not a crawler, see below 200
Applebot-Extended Apple not a crawler, see below 200

Every one returned the same 54,570 bytes a browser gets.

What the check cannot tell you

A 200 from my laptop proves less than it looks like, and I would rather say so than oversell it.

First, I was pretending. The request came from my own address with a borrowed name. It shows that nothing in front of the site blocks by name. It does not show what happens to the real crawler, coming from its own network, if a firewall judges it by address or behaviour. The only full proof is your server log: the real bot, the real status code.

Second, two of my eight names were wrong. Google-Extended and Applebot-Extended are not crawlers. They are words you put in robots.txt to tell Google and Apple whether pages their normal crawlers already fetched may be used for AI. Nothing ever visits your site calling itself Google-Extended, so testing that name with curl tests nothing. My audit of 30 September listed both as crawlers that "receive 200". That line was true and meaningless. For those two, the check is to read your robots.txt, not to send a request.

What the log says

After this post went out I looked at the log I said I had not checked. Cloudflare keeps one for AI crawlers. From 28 September to the morning of 2 October 2026 it counted 959 requests to this site from 17 crawlers, search engines among them, and no crawler was blocked.

Crawler Whose Requests answered
GPTBot OpenAI 217
ClaudeBot Anthropic 158
OAI-SearchBot OpenAI 31
ChatGPT-User OpenAI 27
CCBot Common Crawl 24
PerplexityBot Perplexity 14
Bytespider ByteDance 4
Claude-User Anthropic 3

So the knock on the door was honest: the real crawlers get the same 200 my borrowed names got. This post itself was fetched 17 times in its first three hours. And Google-Extended and Applebot-Extended appear nowhere in the list, which is what you would expect of two names that are not crawlers.

One more thing the log showed. Something calling itself an AI crawler asked this site for files named .env.production and credentials.json. It got a 404 each time, because those files are not there. A crawler's name is only a name. Let the real ones in, and keep nothing on your site that you would mind a stranger asking for.

What was actually blocked: my own images

The audit's first real finding had nothing to do with AI. My robots.txt said:

User-agent: *
Allow: /
Disallow: /og/

It looked tidy: a folder of utility images, kept away from crawlers. But /og/ held en.png and zh.png, the preview image that every page without its own cover hands to anything that asks for one. Preview fetchers that obey robots.txt were being told, by me, not to pick up the picture I had made for them.

The fix was one deleted line. The file is now four lines long, and /og/en.png answers 200 as image/png.

The same pass found three more things of the same kind, each a small lie the site was telling machines:

  • A sitemap address that did not exist. /sitemap.xml returned 404 while Bing still had it on file. It now redirects to the real index.
  • One company, two ids. In the structured data, my profile called BlackRevo by one id and BlackRevo's own page on this site called it by another. To a machine, that is two companies. Now each organisation has one id everywhere, the same one its own website uses.
  • Dates that ran backwards. A post's lastmod in the sitemap was the day it was last saved, in UTC, so a post saved before its publish time claimed to be modified before it existed. It now reads the publish time until a real edit happens.

The Hand fixed all of it on 30 September and left 102 passing tests behind, including ones that pin the ids and the sitemap dates.

What I stopped chasing

Every post here ends with a FAQ, marked up as FAQPage data. I used to think of that as a way to win the expandable questions under a Google result. That prize is gone: Google limited FAQ rich results to government and health sites in August 2023, and its documentation now says they are no longer shown at all from 7 May 2026.

I kept the FAQ anyway and changed what it is for. Each question is now a real heading with its own link, written the way a person would ask it, with an answer that stands on its own. That is not for a rich result. It is for a reader who skims, and for an answer engine that reads a page in pieces and needs each piece to make sense alone.

The heavy thing on this site, the three.js world behind the pages, was measured in the same audit. That became its own story: the night my post pages scored 39.

The five checks

This is the whole method. It takes ten minutes and no paid tool.

  1. Read your robots.txt as a stranger. Open it in a browser. For every Disallow line, name what lives there. If you cannot, find out before a crawler does.
  2. Knock with each crawler's name. Run the curl command above for each real crawler on one page. Same status and same size as a browser, or you have a block to find.
  3. Fetch everything your pages point to. The og:image, the sitemap, the canonical address, the feed. Check the status and the content type, not just that something came back.
  4. One thing, one id. In your structured data, a person or a company gets exactly one id, the same on every page and on every site you control.
  5. Dates that cannot lie. Published, modified and sitemap dates should agree with each other and with what happened.

None of this makes an answer engine quote you. It only removes your own reasons for it not to. What gets quoted is still a page that says one true, specific thing clearly. This list is how you make sure that page can be reached.

FAQ

How do I check if AI crawlers can access my website?

Send a request to one of your pages with the crawler's user-agent string, for example with curl -A "Mozilla/5.0 (compatible; GPTBot/1.1)", and read the status code and size. A 200 with the full page means that name is allowed. Repeat for each crawler you care about, then confirm in your server log that the real bots receive 200 too.

Does Cloudflare block AI crawlers by default?

For domains added since July 2025, Cloudflare blocks known AI crawlers unless the owner chooses to allow them, and it can also manage your robots.txt. On this site, tested on 30 September and 2 October 2026, nothing was blocked and robots.txt was served exactly as written. If your site is on Cloudflare, look at the AI Crawl Control section of the dashboard and run the check yourself.

Is Google-Extended a crawler?

No. Google-Extended is a token for robots.txt that tells Google whether content its normal crawlers fetch may be used for its AI models. Nothing visits your site with that name. Applebot-Extended works the same way for Apple. You control both in robots.txt, not in a firewall.

Is FAQ schema still worth adding in 2026?

Not for a rich result: Google no longer shows FAQ rich results as of 7 May 2026. I still write a FAQ on every post, as real headings with self-contained answers, because readers skim to them and answer engines read pages in pieces. The structured data costs nothing to keep.