What an AI crawler sees when it visits your website
A walk through a made-up bakery website from an AI crawler's point of view, showing what gets read, what gets missed and the simple fixes for each.
· 6 min read
By the Blessed Pixels team
Your website has two kinds of visitor.
One is a person. They see your photos, your colours, the nice font you paid extra for. They scroll, tap and squint at the menu.
The other is a crawler: a piece of software sent out by Google, ChatGPT, Perplexity and the rest to read your site so their AI can talk about you later. It doesn't see colours. It doesn't admire the photos. It reads text, follows links and leaves.
To show the difference, let's follow one around a website. The business is made up, a bakery we'll call Crumb & Co, but every problem below is one we see on real small business sites all the time. If you'd rather skip the story and test your own site, our free AI website check does a version of this walk in under a minute.
7:02am. The front door: robots.txt
Before a well-behaved crawler reads anything, it checks a small file called robots.txt, which sits at the root of every website. This is the house rules: who's allowed in, and where.
Crumb & Co's file has a line a web designer added years ago to "stop spam bots":
User-agent: GPTBot
Disallow: /
That's a closed door for one of OpenAI's crawlers. OpenAI also runs a separate crawler for ChatGPT's search answers, called OAI-SearchBot, and it's listed on their crawler documentation page. Whether to let each one in is a real choice, and some businesses block training crawlers on purpose. The problem is when nobody chose. Plenty of owners have no idea these lines are there.
The fix: open your own robots.txt (your web address followed by /robots.txt) and look for AI crawler names next to Disallow: /. If you didn't decide to block them, ask whoever looks after your site to remove the lines. Our full AI audit covers which crawlers matter and what each one is for.
7:02am and one second. The home page, without the magic
The crawler is in. It asks for the home page.
Here's a surprise for many owners: a lot of AI crawlers read the page as it first arrives, without running the JavaScript that builds parts of modern websites. Vercel looked into this in its write-up The rise of the AI crawler and found that several major AI crawlers, including OpenAI's and Anthropic's, weren't running JavaScript when it checked. Google's crawlers are the main exception.
Crumb & Co's site loads its menu, prices and opening hours with a script after the page appears. People see them. The crawler sees an empty box.
The fix: ask your web person a plain question: "If JavaScript is switched off, is our main text still there?" If the answer is no, the important details need to be in the page itself. This is one of the things we fix most often.
7:02am and two seconds. Reading what's there
What the crawler can read, it reads closely.
The page title says "Home". Not "Crumb & Co, artisan bakery in Bury St Edmunds". Just "Home". Thousands of websites share that title.
The main heading is "Welcome!" Lovely for a person. Useless for a machine trying to work out what this business is.
Then there's a photo of the shop front called IMG_4032.jpg with no description. The crawler can't tell if it's a bakery, a bike shop or a car park.
The fix: give every page a title that says what and where. Make the main heading say what you do ("Sourdough, pastries and celebration cakes in Bury St Edmunds"). Add a short description (alt text) to important images: "Crumb & Co shop front on Abbeygate Street". It takes ten minutes on most website builders, and our free check will tell you if any pages are missing a proper title.
7:02am and three seconds. Looking for the facts
Now the crawler goes hunting for the facts AI tools care most about: name, address, phone, hours.
At Crumb & Co, the address is inside the footer image. The phone number is in a "Contact us" form that never shows the number. The hours are in a nicely designed graphic on the Visit page.
To a person, it's all there. To the crawler, almost none of it is.
Some sites also include a small block of labelled data, called structured data, using the schema.org vocabulary. It spells the facts out for machines: this is a Bakery, here's the address, here are the opening hours. Crumb & Co doesn't have any.
The fix: write your name, address, phone number and hours as plain text on the page, usually in the footer. Then add structured data. Most website builders and WordPress SEO plugins have a section for it, and you can see examples on our what we check page.
7:02am and four seconds. Following the links
The crawler looks for links to other pages and to the business's other profiles. Is this the same Crumb & Co that has a Google Business Profile with 140 reviews? The same one on Instagram?
The website doesn't link to any of them. So the crawler can't be sure.
It also looks for a sitemap, a list of all your pages, and doesn't find one. And it checks for llms.txt, a newer idea proposed in 2024 for giving AI tools a plain summary of a website. It's not a standard yet, and not every AI tool reads it, but it's cheap to add.
The fix: link your website to your Google Business Profile and your main social pages. Make sure your site has a sitemap (most builders make one for you). Add an llms.txt file if your web person can do it easily. You're welcome to look at ours for an idea of the format.
7:02am and five seconds. Gone
That's it. Five seconds and the crawler has moved on to the next site.
What did it learn about Crumb & Co? There's a business. Its home page is called "Home". It says welcome. There might be a photo of something.
Meanwhile the bakery down the road, with half the charm and a quarter of the Instagram followers, has its name, address, hours, specialities and reviews spelled out in plain text. Guess which one an AI assistant feels confident recommending when someone asks for "a good bakery in Bury that does birthday cakes".
This is the uncomfortable truth about AI search. It doesn't reward the best business. It rewards the business it can understand. The good news is that understanding is fixable, usually without a new website. And if your site really is beyond saving, our website plans build all of this in from the start.
Frequently asked questions
Do AI crawlers slow my website down?
On a normal small business site, no. Well-behaved crawlers visit occasionally and read a few pages. If you ever see heavy traffic from one bot, your hosting company can help you limit it without blocking it entirely.
Should I block AI crawlers to protect my content?
It's your choice, and there are fair reasons either way. Some crawlers collect content to train AI models, while others fetch pages to answer a question someone just asked. If you want AI tools to recommend you, blocking the search and answer crawlers works against you.
How often do AI crawlers visit?
It varies by crawler and by site, and the companies don't publish fixed schedules. Busy, frequently updated sites tend to be visited more often. Your hosting logs show which bots have called in.
Is llms.txt worth adding?
It's a low-effort extra, not a must. It's still a proposal rather than an agreed standard, so get the basics right first: readable text, clear titles, contact details and structured data.
How can I see my website the way a crawler does?
Our free AI check reads your site in a similar way and reports what it found. You can also switch off JavaScript in your browser settings and reload your home page. Whatever disappears is what many crawlers miss.