October 1, 2026
I Sent Five AI Crawlers to 29 St. Louis Websites. Seven Slammed the Door.
I expected JavaScript to be the villain. It wasn't. I fetched 29 local business homepages as Chrome, GPTBot, OAI-SearchBot, PerplexityBot, and ClaudeBot, then rendered every one with Firecrawl. 22 let everybody in. Seven had a wall for at least one AI crawler: three Italian restaurants behind a Cloudflare challenge, a dentist that turns away ChatGPT's search bot, a law firm that throttles two of them, and a contractor with an expired certificate. Zero of them asked for any of it in robots.txt.
Everybody says AI crawlers can't read JavaScript. So I went looking for the blank pages.
The standard warning goes like this. Most AI crawlers do not run JavaScript, your fancy site builder renders everything in JavaScript, so ChatGPT sees a blank page and recommends your competitor. It is a tidy story. I tell a version of it myself.
I wanted to know how often it is true for the businesses I actually work around. Not big brands. Dentists, roofers, HVAC shops, injury lawyers, and the Italian places on The Hill. So I picked 29 St. Louis homepages, none of them my clients, and asked each one for its homepage five times, once as each visitor: a normal Chrome browser, GPTBot, OAI-SearchBot, PerplexityBot, and ClaudeBot. One request at a time, a few seconds apart, so nobody could blame the result on me hammering them.
Then I ran every homepage through Firecrawl, which renders the page the way a browser would and hands back clean text. If JavaScript was hiding the business, the gap between the plain fetch and the rendered page would show it.
22 open doors, 7 walls, and a JavaScript gap that barely existed
JavaScript first, because it is the part I was wrong about. Of the 25 homepages that answered a plain request at all, the median page had 1,145 words in its raw HTML and 1,243 after Firecrawl rendered it. Four out of 25 grew by a quarter or more once rendered, and the biggest jumps were partly cookie banners. Fifteen of the 29 run on WordPress, which sends real HTML. The blank JavaScript page was not the problem in this sample.
The walls were. 22 of 29 sites served the same full page to every visitor I sent. The other seven had a door shut on at least one AI crawler.
Three Italian restaurants on The Hill, on two different restaurant platforms, answered every plain request, browser included, with a Cloudflare challenge page titled Just a moment. Zero words. Firecrawl got through and found 713 words on one, and 110 and 157 on the other two, because their homepages are mostly a menu link and an order button. Even rendered, those two never put a phone number or a street address in the homepage text.
One dentist answered GPTBot, OAI-SearchBot, and ClaudeBot with an error and let Chrome and PerplexityBot right in. One personal injury firm told GPTBot and ClaudeBot too many requests on both of my passes while serving everybody else. One concrete company blocked ClaudeBot and nobody else. And one concrete contractor had an expired security certificate, so every plain fetcher I pointed at it failed before it got a single word. Firecrawl read 612 words off the same site.
Last receipt, the one that made me laugh. I checked robots.txt on all 29. Zero of them told GPTBot to stay out. Every one of those seven walls was a firewall, a CDN setting, or a speed plugin making a decision nobody at the business ever wrote down.
Your website has a bouncer. You never met him.
Nobody at that dental office decided to hide from ChatGPT. Something in front of the site decided for them. That is not paranoia, it is policy now: Cloudflare started blocking AI crawlers by default for new customers in July 2025, and security and caching plugins have their own opinions. The business owner sees a normal website in their own browser and assumes everyone else does too.
The bots are not interchangeable, either. OpenAI says on its crawler page that GPTBot gathers training data, while OAI-SearchBot is the one used to surface websites in ChatGPT's search results. You can have a perfectly reasonable opinion about training. Blocking the search bot by accident is a different thing. That is turning away the person who was about to recommend you.
And the boring failures still matter most. An expired certificate does not care what kind of robot you are. A homepage with no phone number and no street address in its text is asking a machine to guess, which is how you end up with the wrong hours and the wrong suite number in an AI answer, the whole subject of the brand-wrong piece. Twenty-six of 29 did put a phone number in plain text. The three that did not are exactly the ones a model would have to piece together from somebody else's directory listing.
Meet your bouncer in fifteen minutes
One. Ask your own homepage who it lets in. From any terminal, curl -A "OAI-SearchBot/1.0" -I followed by your web address, then repeat it with GPTBot, PerplexityBot, and ClaudeBot. You want a 200 every time. A 403, a 400, or a 429 means something in front of your site is answering for you.
Two. If you get a wall, look past robots.txt. In this sample it was never robots.txt. Check the bot settings in Cloudflare or whatever sits in front of the site, then the security and speed plugins. Decide on purpose. Letting the search bots in while keeping the training bots out is a legitimate choice. Blocking all of them because a default said so is not a choice.
Three. Put your phone number and street address in the homepage text. Not only in a button, not only in a logo image, not only in a footer widget that loads later. Words a machine can read without guessing.
Four. Renew the certificate. Set a calendar reminder if your host does not auto-renew it. It is the cheapest fix on this list and it knocked one business out of every plain fetch I made.
Five. Then do the structured part. Sixteen of the 25 readable homepages carried LocalBusiness-style schema. If yours does not, the crawled-not-indexed field guide covers what else the machines need before they take you seriously.
How I ran it, and how you can
The rendering half of this test ran on Firecrawl. You hand it a URL, it loads the page like a browser, and it returns clean markdown, the same shape an AI system wants to read. It got through the Cloudflare challenge pages and the expired certificate that stopped my plain fetches cold, which is why I could tell the difference between a site that has nothing to say and a site that is not letting anyone hear it. The whole test used about 34 credits.
Firecrawl's free tier is 1,000 credits a month with no card, which is enough to check your own homepage and every competitor you care about. If you go past that, signing up through my link gets new users 10 percent off their first purchase. I am not a neutral party here: the AI readiness check inside the app I build for Local Howl runs on Firecrawl too, with a plain fetch as the fallback, because I wanted both views of every client site for the same reason this test needed both.
It's not the JavaScript. It's the door.
Caveats, because I owe you them. One homepage per business, one day, from my office connection. Real AI crawlers come from their own published networks, and a firewall can treat them better or worse than a bot name sent from my desk. Bot rules key on that name, which is what I tested.
But the shape was clear. Most local sites in St. Louis are readable, and the JavaScript scare did not show up. About one in four had a wall in front of at least one AI crawler, and not one of those walls was written down anywhere the owner would ever look. Go meet your bouncer. Then tell him who is on the list.