Your robots.txt has regulated which search engine crawlers can access your website for years. Beyond Googlebot, there are now AI crawlers and search bots that can technically fetch public content. If you do not know these crawlers, you lose transparency over technical accessibility. This guide shows which AI crawlers exist, how to configure them, and where llms.txt only adds context.
GPTBot is one OpenAI crawler. OAI-SearchBot and ChatGPT-User have different roles and should not be mixed with GPTBot. ClaudeBot belongs to Anthropic, PerplexityBot searches the web for Perplexity AI, and Google-Extended is a Google product token for controlling certain uses of crawled content - separate from normal Googlebot. Each crawler has its own user-agent string and can be handled separately in robots.txt. This affects technical access; citations and mentions remain provider decisions.
Try it now
Check your GEO Score in 60 seconds - free, no account needed. 42 factors analyzed.
The configuration follows the same pattern as for Googlebot. With "User-agent: GPTBot" and "Allow: /" you allow technical access. With "Disallow: /internal/" you block specific directories. For machine-readable discoverability, avoid accidentally blocking public product, service and information pages, while keeping admin panels and internal documents blocked. Important: robots.txt controls access, not mentions, citations or recommendations.
Beyond robots.txt, llms.txt is an emerging add-on file, not a universally accepted ranking standard. It sits in your website's root directory and provides a compact summary of your business: who you are, what you offer and what your core products or services are. Systems that evaluate llms.txt receive a structured entry point. Beconova generates your llms.txt automatically from your product data and Schema.org information. This is a useful discovery layer, not a ranking or recommendation promise.
The most common mistake is treating all OpenAI bots as one. OpenAI documents separate user agents: GPTBot collects public content that may be used to train models. OAI-SearchBot indexes pages so they can appear as sources in ChatGPT search features. ChatGPT-User fetches pages when a person in ChatGPT asks for a specific URL or live information. In practice: blocking GPTBot excludes your content from training, but your site stays reachable for ChatGPT search via OAI-SearchBot - unless you block that bot as well. Other providers separate their bots similarly: Anthropic uses ClaudeBot plus dedicated agents for search and user-initiated fetches, Perplexity distinguishes PerplexityBot and Perplexity-User. Google-Extended is not a crawler but a control token: it governs certain uses for Gemini models, and according to Google it does not affect whether your page appears in Google Search. Always verify current names in each provider's official documentation, because user agents can change.
A common setup for businesses that want to be found in AI search: a block "User-agent: OAI-SearchBot" with "Allow: /", a block "User-agent: PerplexityBot" with "Allow: /", and a block "User-agent: GPTBot" with "Allow: /" or "Disallow: /" depending on whether you want to permit training use. Areas such as /admin/, /checkout/ or /account/ are blocked for everyone in a shared "User-agent: *" block. Three rules prevent typical errors: first, each bot follows only the most specific matching block - a bot with its own block ignores the rules under "User-agent: *". Second, a single "Disallow: /" blocks the entire site. Third, the file must live exactly at /robots.txt in the domain root and return status 200. Also reference your sitemap with "Sitemap: https://your-domain.com/sitemap.xml".
Many sites allow AI crawlers in robots.txt and still block them. The most frequent cause is bot protection at CDN or hosting providers: firewalls, bot-defense modes or rules against AI crawlers serve error pages or challenges before robots.txt even matters. A second cause is content that only loads via JavaScript: not every crawler renders JavaScript fully, so key facts belong in the HTML. Third, prices, services or contact details often exist only in images or PDFs - HTML text is far more reliable for machines to read. So check not only the file but your server's actual response: status code, delivered content and load time.
User-agent strings are easy to fake. A request with "GPTBot" in the user agent does not prove that OpenAI actually made it. Several providers publish the IP ranges of their crawlers; OpenAI provides JSON lists for this. Match the IP addresses in your server logs against these lists to separate genuine from spoofed requests. For a quick self-check, Beconova's Robots.txt AI-Crawler Checker shows which known AI user agents your robots.txt allows or blocks. For ongoing detection of real AI visits, Beconova offers AI visitor tracking.
robots.txt and llms.txt are technical foundations for AI-readable content. Configuring both correctly improves accessibility and machine-readable orientation. Whether content is cited or mentioned is decided by the respective provider. Beconova helps with automatic llms.txt generation and the integration guide.
Check GEO Score for freeMarvin Malessa
Founder, Beconova
Founded Beconova in Germany in 2025 to help shops and service businesses become visible in AI search engines. Writes about GEO, AI visibility, and the future of search.
Get started with Beconova now and optimize your presence in AI search engines.
Get Started