Find out who is allowed to read your site
robots.txt and AI crawler checker
Assistants that answer questions with sources can only cite what their crawlers are allowed to read. Enter a domain and this page fetches its robots.txt, applies the rules the way the published specification describes for each of 20 crawler names, and tells you in one sentence whether AI assistants can read your site, with the fixes if they cannot.
Our server requests /robots.txt and /llms.txt, and asks for the headers of /sitemap.xml, /.well-known/ai-plugin.json and /ai.txt. Only public http and https hostnames are requested, each redirect is checked again, and the domain is not stored.
How to use it
- Enter your domain. Type the domain as visitors know it, for example example.com. A full address works too and is reduced to its domain.
- Run the check. Press Check site. The files are fetched once and the rules are applied for every listed crawler.
- Read the verdict. The top line says whether AI assistants can read your site, only some can, or none can, based on the homepage.
- Scan the table. Each row shows one user agent against the homepage, /tools and /pricing. Hover or tap a cell for the rule that decided it.
- Apply the fixes. Edit robots.txt, add or repair llms.txt, declare your sitemap, then run the check again to confirm.
How the rules are applied
The robots.txt matcher follows RFC 9309 (the Robots Exclusion Protocol) and Google's documented interpretation of it (developers.google.com/search/docs/crawling-indexing/robots/robots_txt). User agent lines open a group, and a crawler uses the group or groups that name its token exactly, ignoring case, and otherwise the group for the star user agent. Groups that name the same token are merged. Rules in a named group replace the star group completely.
Within the chosen group, a rule matches when its path pattern matches the start of the URL path. A star matches any run of characters and a dollar sign at the end anchors the pattern to the end of the path. The most specific rule wins, meaning the longest pattern, and when an allow and a disallow have the same length the allow wins. An empty Disallow matches nothing. Paths are case sensitive.
A robots.txt that answers with a 4xx status is treated as no restrictions, and one that answers with a server error or cannot be reached is reported as unavailable, which crawlers may treat as a full block. The llms.txt check looks for Markdown that starts with an H1 and contains links. A robots.txt rule is a request that well-behaved crawlers honour, not access control, and the crawler names are listed as data: names and rules change, so confirm against each operator's own documentation.
Questions
Does blocking a crawler in robots.txt remove my site from its answers?
It asks that crawler not to fetch the pages. It does not remove anything already collected, and crawlers that ignore the file are unaffected. Treat robots.txt as a polite request, not a lock.
Why are some crawlers allowed only because of the star group?
When no group names a crawler, it falls back to the star group. The table shows which group was used in the tooltip of each cell, so you can see whether a rule was written for that crawler or inherited.
What does a missing robots.txt mean?
A 404 means there are no restrictions, so every crawler is allowed. That is a valid setup, but you lose the place to declare your sitemap.
What should llms.txt contain?
A Markdown file at /llms.txt that starts with an H1 naming the site, ideally a short blockquote summary, then sections of links to the pages you most want read. This tool checks that shape, not whether any assistant uses the file.
Are ai-plugin.json and ai.txt required?
No. They are older or informal conventions, and the tool only reports whether the files exist. Nothing here depends on them.
Why can some domains not be checked?
Only public hostnames are fetched. A name that points at a private network, an IP address, a custom port or a login in the address is rejected before any request, as is a redirect that leads there. Those sites cannot be checked from here.