🤖 Robots.txt Checker
View and analyze the robots.txt file for any website to see crawler directives.
What this checker shows you
Enter a domain and this tool fetches /robots.txt from that
site and prints the file back to you exactly as it was served. It is a
reader, not a validator: nothing is interpreted, corrected or scored, so
what you see is what a crawler sees.
That matters more than it sounds. A robots.txt file is plain text with no error reporting. Nothing warns you when a rule is wrong — the file is simply obeyed, and traffic quietly stops arriving.
How to read the file
Rules come in groups. A group starts with one or more
User-agent lines naming the crawlers it applies to, followed
by the rules for them.
User-agent: *
Disallow: /admin/
Allow: /admin/public-notice.html
Sitemap: https://example.com/sitemap.xml
| Directive | What it does |
|---|---|
User-agent | Names the crawler the group applies to. * means every crawler that has no group of its own. |
Disallow | A path prefix the crawler should not request. Disallow: / means the whole site. |
Allow | Carves an exception out of a broader Disallow. The more specific rule wins, not the later one. |
Sitemap | An absolute URL to a sitemap. It is not tied to a group and can appear anywhere in the file. |
Crawl-delay | Seconds to wait between requests. Bing and Yandex honour it; Google ignores it entirely. |
An empty file, or no file at all, means everything may be crawled. So does a file that returns a 404 — which is why a missing robots.txt is not an error and does not need fixing.
Mistakes worth looking for
A Disallow: / that came from staging
The single most damaging line in SEO. Staging sites are blocked wholesale, then the file ships to production with the launch. Traffic does not fall immediately — it decays over weeks as pages drop out of the index, which is what makes it so easy to miss.
Expecting robots.txt to hide a page from search
It does not. Disallow stops a crawler fetching a
page, not listing it. A blocked URL that other sites link to can
still appear in results, showing the bare URL with no description,
because Google knows it exists but was never allowed to read it.
<meta name="robots" content="noindex"> or an
X-Robots-Tag header. Blocking it in robots.txt guarantees the
noindex is never seen.Blocking the CSS and JavaScript
Common in older files that disallow /assets/ or
/static/ wholesale. Google renders pages before judging them;
a page whose stylesheet it cannot fetch is assessed on a broken layout,
and mobile-friendliness suffers with it.
The file is not where you think
robots.txt is read only from the root of a host, and each host is
separate. example.com, www.example.com and
shop.example.com each need their own file, and
https:// is distinct from http://. A file at
/blog/robots.txt is never consulted by anything.
Paths are case-sensitive
Disallow: /Admin/ does not block /admin/.
Directive names are not case-sensitive, but the paths after them are.
Frequently asked questions
Does a missing robots.txt hurt my site?
No. If the file is absent the site is fully crawlable, which is the normal state for most sites. Add one when you have something specific to exclude, or to point crawlers at your sitemap.
How soon does a change take effect?
Google caches robots.txt for around 24 hours, so a fix is not instant. In Search Console the robots.txt report shows the version Google is currently working from and lets you ask for a re-fetch.
Do all crawlers obey it?
The major search engines do. It is a convention, not a security
control — a scraper that ignores it faces no obstacle. Anything that
must stay private needs authentication, not a Disallow.
Can I use wildcards?
Google, Bing and Yandex support * for any sequence of
characters and $ to anchor the end of a URL, so
Disallow: /*.pdf$ blocks PDFs. Smaller crawlers may treat
those characters literally, so do not rely on them for rules that
matter.
Why does the tool show nothing for my domain?
Either the site has no robots.txt, or it did not answer. Both are
reported the same way here. Open https://yourdomain/robots.txt
in a browser to tell the two apart.