What it looks like
Robots.txt is a plain text file at the root of a site, for example example.com/robots.txt. A small one looks like this:
User-agent: *
Disallow: /cart/
Disallow: /search/
Sitemap: https://www.example.com/sitemap.xml
The first line says the rules are for all crawlers. Each Disallow line names a part of the site they should not visit. The Sitemap line is optional and tells crawlers where to find your list of pages.
Why it matters
- It keeps crawlers out of areas that waste their time. Cart pages and internal search results are common examples. On very large sites this protects crawl budget. Small sites rarely need to think about it.
- It controls crawling, not indexing. Crawling is the visit. A blocked address can still appear in search results as a bare link if other pages link to it.
- One line can block everything.
Disallow: /tells crawlers to stay away from the whole site. - It is not a lock. The file is public and anyone can read it. It does not protect private pages.
COMMON MISTAKELeaving a rule that blocks the whole site in place after launch. Sites are often blocked while they are being built. If the rule stays, crawlers keep away from every page.
How to check yours
- Type your domain followed by
/robots.txtinto your browser. If no file exists, crawlers treat the whole site as open. - Read each
Disallowline. Make sure none of them covers a page you want in search. - Look for a
Sitemapline that points to your XML sitemap. - To keep a page out of the results, do not block it here. Leave it open and use noindex instead.
The difference between the two is the heart of the lesson crawling and indexing, explained simply.