How it works
Search engines use programs called crawlers to visit pages. Google's crawler is called Googlebot. A crawler does three things, over and over.
- It finds an address. Mostly by following links from pages it already knows, and by reading an XML sitemap.
- It checks the rules. Before visiting, it reads the site's robots.txt file to see which addresses it should stay away from.
- It fetches the page. It downloads the page, reads the content and notes every link it finds. Those links go on its list of addresses to visit next.
Crawling comes before indexing. Crawling is reading the page. Indexing is deciding to store it. A page can be crawled and still not be indexed.
An example
You publish a new article and link to it from your homepage. The next time Googlebot visits the homepage, it sees the new link and adds the address to its list. Some time later it visits the article and reads it. That visit is the crawl. It may happen within days or take a few weeks, and nobody can promise a date.
Why it matters
- A page that is never crawled cannot be indexed, so it cannot rank.
- Pages with no links pointing to them are hard for crawlers to find.
- A wrong rule in robots.txt can keep crawlers away from pages you want in search.
- Crawlers return to pages they know, which is how your changes get picked up.
COMMON MISTAKETreating a blocked page as a hidden page. Blocking crawling in robots.txt does not reliably keep an address out of search results. To keep a page out, allow crawling and add a noindex tag.
How to check yours
- Open Google Search Console and paste the page's full address into the URL inspection bar.
- Read the last crawl date and whether crawling was allowed.
- If the page has not been crawled, add an internal link to it from a page that is already in Google.
- Read crawling and indexing, explained simply for the full sequence.