Making Your Site Readable to AI: A Technical Guide to Crawlers, Structure and Access
The boring layer underneath everything else, where the most expensive mistakes live because they fail silently.
Most conversations about AI search jump straight to content. Write clearer answers, publish original data, get more reviews. All good advice, and all of it useless if the systems doing the reading cannot reach your pages in the first place.
This is the boring layer underneath everything else. It is also the layer where the most common and most expensive mistakes live, because they are silent. Nothing breaks. No error appears in your reporting. Your site simply stops being an option.
Here is what to check and how to fix it.
Two different jobs, two different bots
The first thing to understand is that AI systems visit your site for two distinct reasons, and you may want to treat them differently.
Training crawlers collect content to be used in building or improving models. What they take shows up in a model's general knowledge, without a link back to you, and it appears months later.
Live retrieval fetchers visit at the moment a user asks a question. They pull your page, extract the relevant part, and the assistant may cite you directly in the answer with a link.
That difference matters enormously for policy. Blocking training crawlers is a legitimate business decision that many publishers have made. Blocking live retrieval fetchers means opting out of being cited in answers, which for most commercial sites is self defeating.
The problem is that plenty of sites block both without realising, usually because someone applied a blanket rule, or because a security vendor did it for them. If you take one thing from this article, make it this: go and check what your site actually allows, today, rather than what you think it allows.
Check your robots.txt properly
Open yourdomain.com/robots.txt and read the whole thing, including inherited rules and wildcards.
Look for blanket disallow rules affecting all user agents. Look for rules targeting AI specific user agents by name. Look for rules blocking directories where your most quotable content lives, like /docs, /blog, /pricing or /resources. Staging leftovers are a classic problem, where somebody blocked everything during a rebuild and nobody removed it afterwards.
Be aware that robots.txt is voluntary. Well behaved crawlers respect it. It is not a security mechanism, and if you genuinely need to prevent access you need server side controls.
If you decide to allow retrieval while restricting training, write the rules explicitly per user agent rather than relying on a general allow with exceptions. Explicit rules are easier for the next person to understand and harder to get wrong.
Then check the actual behaviour, not just the file. Look at your server logs and confirm that AI user agents are fetching pages and receiving 200 responses. This is the only way to know for sure.
Your CDN or WAF may be blocking bots you never blocked
This is the single most common hidden cause of invisibility, and it usually has nothing to do with your marketing team.
Content delivery networks and web application firewalls offer bot management features, and several have shipped settings that block AI crawlers by default or as part of a recommended security preset. Somebody in engineering ticked a box during a security review, and now your entire site is unreachable to the systems that decide whether you get mentioned.
Check your CDN's bot management rules. Check rate limiting, since aggressive limits can throttle crawlers into failure even without an outright block. Check geographic restrictions, because fetches often originate from data centre ranges in regions you may have restricted. Check whether JavaScript challenges or interstitial security checks are being served to non browser clients, which will fail every time.
Ask your infrastructure team directly: what is our current policy for AI user agents at the edge? If nobody knows the answer, that is your answer.
Render on the server
Retrieval systems fetch pages under time pressure. Some execute JavaScript, many do not, and none of them will wait around while your single page application hydrates.
If the meaningful content of your page is injected client side, you are gambling. The safe approach is server side rendering or static generation for anything you want quoted. Product pages, pricing, documentation, blog content, comparison pages, FAQ content.
Test it the direct way. Fetch your page with curl and no JavaScript execution, then read the raw HTML that comes back. If your pricing table, your feature list or your article body is missing from that output, it does not exist as far as a large number of retrieval systems are concerned.
Pay particular attention to content loaded on interaction. Accordions, tabs and "read more" expanders are fine if the text is present in the HTML and merely hidden by CSS. They are a problem if the text is fetched only when the user clicks.
Use HTML structure the way it was designed
Retrieval systems chunk documents, and they use your HTML structure to decide where the boundaries fall. Clean semantic markup produces clean chunks. Div soup produces chunks that start mid sentence and end mid list.
Practical rules:
One H1 per page, describing what the page is about.
H2 and H3 in real hierarchical order, with no level skipping and no headings chosen because of how they look.
Real list markup for lists, real table markup for tabular data. A table built from divs is not a table to a parser.
Paragraphs as paragraphs, not line breaks inside one giant block.
Keep navigation, cookie banners, promotional interstitials and footer content out of the main content flow. Use the main element to mark your primary content. This helps extractors separate what your page is about from what your page also contains.
Also worth checking: whether your key pages have enormous amounts of boilerplate before the content starts. If the first 2,000 characters of your HTML are menus and scripts, some extractors will do a worse job.
Structured data still earns its place
Schema markup does not guarantee anything, and the current generation of language models reads plain text perfectly well. It still helps, for a simple reason. Structured data states facts unambiguously in a machine readable format, which reduces the chance of misinterpretation.
The types worth implementing:
Organization, with your legal name, alternate names, logo, founding date, location and sameAs links to your official profiles. This is your identity record and it helps systems connect mentions of your brand across the web.
Product and Offer, with current pricing and currency. If you want assistants to state your pricing correctly, publishing it in structured form is the most direct thing you can do.
FAQPage for genuine question and answer content.
Article, with author, datePublished and dateModified. The modified date is the one that matters for freshness.
BreadcrumbList, for context about where a page sits.
Two rules. Keep schema in sync with what is visible on the page, since contradictions are worse than no markup at all. And validate it, because broken schema is common and silent.
The llms.txt question
You will have seen suggestions to add an llms.txt file to your root, a markdown document that summarises your site and points to your most important pages in a clean, easily parsed format.
An honest assessment: it is cheap to create, it does no harm, and support for it is not universal. Treat it as a low cost bet rather than a core tactic. If you have a documentation heavy site, it is more likely to be worth the effort, because the format suits that content well.
If you do publish one, keep it short, keep it accurate, list your genuinely most useful pages with one line descriptions, and update it when your site changes. A stale llms.txt pointing at dead URLs is worse than none.
Do not let it distract from the things that definitely matter, which are access, rendering and content quality.
Consistency across your own site
Retrieval systems build confidence from corroboration. When your own site contradicts itself, you undermine that.
Audit for conflicting facts. Pricing that differs between your pricing page, a landing page and an old blog post. Feature lists that describe a product you have since changed. Company details, team size, office locations and certifications that were accurate three years ago.
Check duplicate content too. If the same content is reachable at multiple URLs, with and without trailing slashes, on multiple subdomains or in multiple language versions without proper hreflang, you split the signal and increase the chance of an outdated version being the one retrieved.
Set canonical tags properly. Redirect old URLs rather than leaving them live. Delete or clearly mark obsolete content rather than leaving it to be found.
Files, formats and the things that hide text
A few recurring problems worth a specific check.
PDFs. Frequently used for whitepapers, price lists and technical specs. Parsing is inconsistent. If a PDF contains information you want quoted, publish an HTML version of the same content.
Images containing text. Pricing tables, comparison charts, statistics inside infographics, quotes rendered as graphics. Every one of these is invisible text. Put the same information in HTML alongside the image and write real alt text.
Video and audio. Publish transcripts. A 40 minute podcast episode where you explain something in detail is worthless for retrieval without a transcript, and highly valuable with one.
Gated content. Anything behind a form does not exist. If a resource is important to your positioning, consider publishing a substantial ungated version and gating something else.
Speed, uptime and the unglamorous basics
Retrieval happens live and under a deadline. A page that takes six seconds to respond may simply be dropped in favour of a competitor's page that responds in 400 milliseconds.
Keep server response times low, particularly for the pages you most want cited. Avoid redirect chains, since each hop costs time and some fetchers follow only a limited number. Make sure your TLS certificate is valid and your configuration is current, because certificate errors cause silent failures. Monitor uptime, remembering that a fetch failing during an outage is a citation opportunity lost permanently, not one that retries later.
Serve consistent content to all clients. If you vary your response based on user agent in ways that show bots something different from users, you are creating problems for yourself.
A checklist you can run this week
Fetch and read your robots.txt in full, checking for blanket and AI specific rules.
Ask your infrastructure team what your CDN or WAF does with AI user agents, and check the bot management settings yourself.
Grep your server logs for AI user agents and confirm they are receiving 200 responses.
Curl your five most important pages with JavaScript disabled and confirm the main content is present in the raw HTML.
Validate structured data on your homepage, pricing page and top three content pages.
Search your site for pricing and feature claims, and reconcile any contradictions.
Identify any critical information that exists only inside images, PDFs or video, and publish HTML versions.
Check response times and redirect chains on your commercially important URLs.
Decide your position on training crawlers versus retrieval fetchers, write it down, and make sure the implementation matches the decision.
None of this is exciting work. It is also the difference between a content programme that produces citations and one that produces nothing at all, and it takes a few days rather than a few quarters. Do it before you write another word.