FREE TOOL

Robots.txt tester

Paste your robots.txt and the URLs you care about. See whether each one is allowed or blocked, and which rule decided.

Rules are matched the way Google documents them: the crawler's own group wins over the * group, the longest matching rule wins, and Allow beats Disallow on a tie. Only the path and query of a URL are tested; the host is ignored.
THE RESULT FOR GOOGLEBOT
  • Allowed

    https://example.com/

    Decided by Allow: / in the * group (there is no group for this crawler).

  • Blocked

    https://example.com/admin/users

    Decided by Disallow: /admin/ in the * group (there is no group for this crawler).

  • Blocked

    https://example.com/blog/post?sort=new

    Decided by Disallow: /*?sort= in the * group (there is no group for this crawler).

  • Allowed

    https://example.com/private-images/logo.png

    Decided by Allow: / in the * group (there is no group for this crawler).

WHAT THE FILE SAYS

  • **: 3 Disallow, 1 Allow
  • Ggooglebot-image: 1 Disallow, 0 Allow
  • Ggptbot: 1 Disallow, 0 Allow

Sitemap: https://example.com/sitemap.xml

CHECK THE LIVE PAGE

Paste a URL instead and we will read the real robots.txt and 29 more things about the page.

Run the SEO audit

How to read the result

A crawler reads robots.txt before it fetches a page. It looks for the group whose User-agent line names it; if there is none, it uses the * group; if there is no * group either, nothing is blocked. Inside the group, every Allow and Disallow line is compared with the path of the URL, and the longest matching rule wins. When an Allow and a Disallow match with the same length, Allow wins. This tool applies exactly those rules, which are the ones Google documents, and shows you the line that decided.

Only the path and query of a URL are tested, from the first slash. A pattern can contain * for any run of characters and end in $ to mean the end of the URL, so Disallow: /*?sort= blocks every address with ?sort= in it, and Disallow: /*.pdf$ blocks addresses that end in .pdf. A Disallow with nothing after it allows everything, and a line starting with # is a comment.

Blocking a page in robots.txt keeps the crawler from fetching it; it does not keep the address out of Google. A blocked page can still be listed if other pages link to it, just without a description. To keep a page out of the results, let the crawler fetch it and use a noindex tag instead.

Common mistakes this catches

A Disallow: / left over from a staging site blocks everything. A group for a specific crawler replaces the * group for that crawler, so User-agent: Googlebot followed only by Disallow: /private/ means Googlebot ignores every rule in your * group. And a rule such as Disallow: /blog also blocks /blog-tips and /blogroll, because rules match the start of the path; write /blog/ or /blog$ when you mean the one page.

LEARN MORE

The lessons behind this tool.

TECHNICAL SEO · 01How Google finds your pages.

Crawling and indexing, explained simply

How Google discovers, crawls and indexes a page, why robots.txt and noindex are different, and how to check if your page is indexed.

0114 Sept5 minRead tip
TECHNICAL SEO · 02Your sitemap is a list of promises.

XML sitemaps: what to include and what to leave out

What belongs in an XML sitemap, what to leave out, how to keep lastmod honest, and how to tell Google where the file is.

0224 Sept5 minRead tip