LLMs.txt & Robots.txt: Optimizing for AI Bots | Goodie

LLMs.txt & Robots.txt: Optimizing for AI Bots & Answer Engines

Optimizing for AI bots using robots.txt & LLMs.txt boosts visibility in ChatGPT, Gemini, and answer engines. Learn key crawlability best practices.

by: Rebecca Gross
Published: January 28, 2026

Advanced large language models (LLMs) now directly synthesize answers, often without the user clicking any “blue links.” This has changed the nature of search behavior as a whole: Google SERPs are now dominated by AI Overview (for 60% of searches, in fact), and users are increasingly turning directly to LLMs for answers or information rather than clicking links until they find what they seek. To be exact, 60% of Google’s searches are zero clicks in 2026. This shows us that users are finding their answers directly in the Google AI Overviews.

Because of this change in search behavior, simply ranking #1 in Google’s search results using traditional SEO methods no longer guarantees visibility (especially if AI can serve up the answer directly within seconds).

One key to Answer Engine Optimization (AEO) is making your website AI-friendly. That means understanding how LLM crawlers (the bots that feed content to AI models) differ from traditional search engine crawlers, and how to use files like robots.txt and the emerging LLMs.txt to your advantage.

In this guide, we’ll break down the differences between AI and search crawlers, tackle crawlability challenges (like rendering and timeouts), outline best practices to keep your content in top shape for AI bots, and provide a detailed look at LLMs.txt and robots.txt, including how they work together.

AI Crawlers: LLM Bots vs. Traditional Search Bots

Before diving into the specifics, it’s important to understand what LLM crawlers are and how they operate differently from Google’s or Bing’s usual bots. Traditional search engine crawlers (e.g. Googlebot, Bingbot) scour the web to index content for search results, building an index so they can rank pages on a SERP.

In contrast, LLM crawlers crawl webpages to gather information for an AI model; either to train the model’s knowledge base, or to fetch up-to-date info for answering queries. In essence, they help create a library of resources that LLMs rely on to answer user questions instead of just generating a list of search links.

The differences between LLM bots and traditional search bots can be broken down into three key categories: Purpose, Behavior, and Impact on Traffic.

Purpose

Search bots index pages to later retrieve links for queries. LLM bots, on the other hand, crawl pages to supply AI systems with content that might be later used within an answer.

For example, when an answer engine like Perplexity responds to a question, it may have already crawled and ingested content from several sites, which it then quotes and summarizes on demand.

Examples of LLM Crawlers: Each major LLM or AI search platform has its own bot. OpenAI’s GPTBot and Anthropic’s Claude Crawler are two prominent ones. Other examples include Google-Extended (used by Google and other AI systems for training data) and many more research or startup AI crawlers.

To put the power of these LLM crawlers into perspective, when put together, requests generated by GPTBot and Claude in just one month of late 2024 made up about 20% of Googlebot’s requests from the same timeframe.

These bots identify themselves via User-Agent strings (e.g. GPTBot for OpenAI, Google-Extended for Google’s AI crawler, etc.), which we can target in robots.txt.

Behavior

Traditional crawlers like Googlebot have become very advanced: they execute JavaScript, respect crawl budgets, and avoid overloading sites. Many AI bots are newer and less sophisticated in their crawling behavior. They may not render client-side scripts, and some operate with shorter timeouts or simpler link discovery.

This means that it is imperative to serve as much pertinent info in HTML as possible. If there is JavaScript on the page, LLM bots will likely not crawl it. In addition to this, if important info lies underneath JavaScript (such as information hidden in a dropdown), it will likely not be crawled even though it is in HTML.

In short, structure your most important info at the top of the page simply and smartly.

AI crawlers often perform two roles: one set of bots gathers broad data for training the model, and another set (for some platforms, like ChatGPT’s search feature) does real-time crawling for Retrieval-Augmented Generation (RAG), which allows it to pull in fresh data when the AI needs up-to-the-minute info.

For example, an LLM might use a real-time crawler to fetch today’s news while relying on its training index for older knowledge.

Impact on Traffic

Don’t be surprised to see AI bots in your traffic logs. As noted, OpenAI and Anthropic’s bots are already crawling at significant scale. Unlike human users, these bots won’t show up in analytics dashboards as pageviews, but they consume bandwidth.

If you run a popular site, a wave of AI crawlers could hit your pages. The upside is increased chances of your content being used in AI answers; the downside is potential load on your servers if not managed well.

LLMs.txt: A New File for AI Crawlers (What It Is & How to Use It)

You’re probably familiar with robots.txt: a file that tells crawlers what they can’t access. LLMs.txt is a newer concept often described as a robots.txt for AI models. But instead of blocking bots, LLMs.txt is about feeding content directly to them in a structured way.

What Is LLMs.txt?

LLMs.txt is essentially a special text (Markdown) file at your site’s root (yourdomain.com/llms.txt) that provides a curated guide to your site’s important content, specifically for Large Language Models. The idea was proposed in 2023 as LLMs started becoming mainstream.

Think of it as a cheat sheet for AI: it points to your most valuable pages (especially things like documentation, FAQs, product info, and policies) in a simplified format that an LLM can easily consume.

Unlike robots.txt, LLMs.txt doesn’t use “Disallow” rules or tell bots what not to do. Instead, it highlights what content the AI should focus on. For example, you might use LLMs.txt to say, “Hey AI, here’s a quick overview of our site and links to our key resources.” This can help the model understand your content without wading through unnecessary stuff like navigation menus or ads.

LLMs.txt Format & Structure

The LLMs.txt file is typically written in Markdown, which is both human-readable and easy for machines to parse. There’s a suggested structure that has become common:

For example, an LLMs.txt for a company might look like:

In this hypothetical snippet, the website owner has created Markdown versions of key pages (e.g. the API page and the Features page) and listed them. The LLMs.txt provides an at-a-glance map of the site’s crucial info. An AI crawler could fetch this single file and quickly know where the “good stuff” is, rather than guess by crawling every page.

The goal is to remove ambiguity by giving AI a prioritized content list.

General Guidelines for LLMs.txt:

Now, it’s important to set expectations. LLMs.txt is not yet widely adopted by the major AI players. OpenAI’s GPTBot, Google’s crawlers, and others do not currently check for or use LLMs.txt by default.

Google’s own search representatives have even likened LLMs.txt to an early, unused idea: the old meta keywords tag in terms of current impact. Many webmasters who implemented LLMs.txt reported that known AI bots never requested the file at all. In one case, thousands of domains saw virtually zero hits on LLMs.txt except from a few curious minor bots.

Ensuring AI Crawlability: Key Challenges & Solutions

Now that we’ve covered what LLMs.txt is and how to create it for your site, we’ll cover some other ways that you can ensure AI crawlers can “read” your site. This is the foundation of answer engine optimization, or AEO.

JavaScript & Rendering

As we touched on before, most AI crawlers do not execute JavaScript. Google’s crawler can render JS, but bots like ChatGPT’s or Claude’s essentially only fetch raw HTML. In fact, OpenAI and Anthropic’s bots still attempt to fetch some .js files, but they don’t run them. In one analysis, ~11.5% of ChatGPT’s requests were JS files (23.8% for Claude) that likely went unused.

Solution: Use server-side rendering (SSR) for key pages or provide a static HTML fallback. Deliver the core text in the initial HTML response so the bot sees it without needing JS. You can still use JS for interactive widgets, but critical text (product descriptions, article content, etc.) should be present in the HTML source. This way, even a less sophisticated AI crawler can read it.

Page Load Speed & Timeouts

AI crawlers can be impatient. Many operate with very short timeouts (often 1-5 seconds) when fetching content. If your webpage is slow or heavy, the bot might give up or only be able to grab part of the content before it times out.

Solution: Optimize your site performance. Compress images, use efficient code, and consider caching to speed up delivery. Aim for a sub-2 or 3 second time-to-first-byte (TTFB) to account for LLM bots.

On a similar note, surfacing important info high in the HTML is helpful. If an AI bot times out after a few seconds, you want it to have seen your key points or introductory paragraphs. In practice, this means having a clean (not loading 50 scripts before content) and maybe a summary or intro at the top of your page that quickly communicates the page’s main topic.

HTML Structure & Semantics

AI models are better at parsing content when it’s well-structured. They don’t have the millions of human-like judgment calls that Google’s algorithm has developed over decades of evolution; they often just consume text as-is.

Solution: Use clean, semantic HTML. Proper heading hierarchy (

…

…) and clearly labeled sections help the crawler (and the AI) understand the context. For example, use

for subheadings,