TECHNICAL SEO · 01
How Google finds your pages.

Crawling and indexing, explained simply

How Google discovers, crawls and indexes a page, why robots.txt and noindex are different, and how to check if your page is indexed.

You publish a page, search for it a week later, and it is nowhere. Before you change anything, it helps to know how a page gets into Google in the first place. This lesson explains the steps in plain terms and shows you how to check where your page got stuck.

The three steps: crawl, index, rank

Every page goes through the same three steps.

  1. Discover and crawl. Google first has to learn that the address exists. Then its program, Googlebot, visits the page and reads it. That visit is called crawling.
  2. Index. Google processes what it read and decides whether to store the page in its index, which is a very large library of pages. This is indexing. Not every crawled page gets indexed.
  3. Rank. When someone searches, Google picks pages from the index and puts them in order.

The order matters. A page cannot rank if it is not indexed, and Google cannot judge a page it has not crawled. So when a page is missing, check the steps in this order and stop at the first one that fails.

How a page gets discovered

Google finds new pages in two main ways.

Use both. A sitemap tells Google that a page exists. A link does that too, and it also shows that the page matters to you.

Discovery takes time. For a new page it can be a few days or a few weeks, and nobody can promise a date. A new site with few links pointing to it usually waits longer.

Blocking crawling is not blocking indexing

There are two common ways to keep Google away from a page. They do different things, and mixing them up is the most common technical mistake I see.

robots.txt is a text file at the root of your site. It tells crawlers which addresses they should not visit. Here is a small robots.txt file:

User-agent: *
Disallow: /cart/
Disallow: /search/

It says: all crawlers, do not visit any address that starts with /cart/ or /search/.

noindex is an instruction on the page itself. A noindex tag says: you may read this page, but do not show it in search results. It goes in the head of the page:

<meta name="robots" content="noindex">

The difference matters for two reasons.

First, robots.txt controls crawling, not indexing. If a blocked page is linked from somewhere else, Google can still index the address without reading the page. It can then show up in results as a bare link with no description.

Second, a noindex tag only works if the page can be crawled. If robots.txt blocks the page, Google never reads the tag.

BEFORETo keep a page out of Google, block it in robots.txt.
AFTERTo keep a page out of Google, allow crawling and add a noindex tag.

Google cannot follow an instruction on a page it is not allowed to read.

So which one should you use?

  • Use robots.txt for areas where crawling is a waste of time, such as cart pages and internal search results. On very large sites this protects crawl budget, the number of pages Google is willing to crawl on your site in a given period. Small sites rarely need to think about it.
  • Use noindex for pages people may visit but that should not appear in search, such as a thank-you page after a form.
  • Use a login for anything private. Neither robots.txt nor noindex protects a page from visitors.
COMMON MISTAKEDo not block a page in robots.txt and add noindex to it as well. The block hides the noindex tag, so the address can stay in the results.

How to check if a page is indexed

The most reliable check is the URL inspection in Google Search Console, a free tool from Google for site owners. After you verify your site, paste the full address of a page into the inspection bar at the top. The report tells you:

  • whether the address is on Google
  • how Google discovered it, including sitemaps and pages that link to it
  • when it was last crawled
  • whether crawling and indexing were allowed
  • which address Google chose as the main version

You can also test the live page and ask Google to index it. That is a request, not a guarantee.

For a quick look without Search Console, search for site: followed by the page's address, with no space after the colon. It is a rough check, not an exact one.

Common reasons a page is missing

  • Nothing links to it. Google has not found it yet.
  • robots.txt blocks it. Often a rule left over from when the site was built.
  • It has a noindex tag. Often left on from a test version, or switched on by a "discourage search engines" setting.
  • The canonical points elsewhere. A canonical tag is a hint about which address is the main version of a page. If it names another page, Google may index that one instead.
  • It redirects or returns an error. Google indexes the destination, or nothing.
  • It is very similar to another page. Google may keep only one of them.
  • It is new. Sometimes the answer is to wait.

This is the first lesson in the Technical SEO series.

Check your own site

Pick one page that should be in Google and go through it:

  • Does at least one other page on your site link to it?
  • Is it listed in your sitemap?
  • Does robots.txt allow it to be crawled?
  • Is it free of a noindex tag it should not have?
  • What does the URL inspection say about it?

For a wider look at the same page, run the free audit.

That’s lesson 1. Want me to check your own page?

Audit my page
Ona

I work on the technical side of SEO. SEO Expert Tips is where I break that work down into small, practical steps. More about me