Many sitemaps are generated once and never looked at again. Over time they fill up with redirects, duplicates and pages nobody wants in search. This lesson shows you what belongs in a sitemap, what to leave out, and how to hand the file to Google.
What a sitemap is for
An XML sitemap is a file that lists the addresses on your site that you want search engines to find. It helps with discovery, the first step of crawling. It is most useful when a site is new and has few links pointing to it, when a site is large, or when some pages sit many clicks away from the homepage.
A sitemap does not guarantee that a page gets indexed, and it does not improve rankings. It is a list of suggestions. If the difference between found and indexed is new to you, read crawling and indexing, explained simply first.
I think of a sitemap as a list of promises. Every address in it promises three things: this page exists, this is the version I want shown, and it is worth showing. The cleaner the list, the more useful it is.
Here is a minimal sitemap with two pages:
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://www.example.com/</loc>
<lastmod>2026-09-12</lastmod>
</url>
<url>
<loc>https://www.example.com/services/roof-repair</loc>
<lastmod>2026-08-30</lastmod>
</url>
</urlset>
Each entry needs only one thing: loc, the full address of the page. lastmod is optional. Older sitemaps often include priority and changefreq as well. Google ignores both, so you can leave them out.
What belongs and what does not
Include an address only if all three of these are true:
- It loads normally. The server answers with status 200, which means "here is the page".
- It can be indexed. It has no noindex tag and is not blocked in robots.txt.
- It is the canonical version. A canonical tag tells search engines which address is the main version when several show the same content. It is a hint, and the sitemap is a second hint. The two should agree.
Leave these out:
- Redirects. A redirect sends visitors on to another address. List the final address, not the old one.
- Error pages. Anything that returns "not found" or a server error.
- Pages with noindex. A noindex page in a sitemap says "find this" and "do not show this" at once.
- Duplicates. The same page with and without a trailing slash, with tracking codes, or on both http and https.
- Filtered, sorted and search result addresses. A shop can generate thousands of these from one category page.
- Utility pages. Cart, login, thank-you pages.
The shop example looks like this:
The first address is one of many views of the same category. The second is the page you want people to land on. Only the second belongs in the sitemap.
A short, clean list is better than a list of every address your site can produce.
Keep lastmod honest
lastmod is the date a page last changed in a way that matters. Google uses it only when it is consistently accurate. The usual way to break it is a system that stamps every page with today's date each time the sitemap is rebuilt. If every date changes every day while the pages stay the same, the dates say nothing, and Google has no reason to rely on them.
A change that matters is a rewritten section, new prices, a new product, an updated opening time. A new year in the footer is not.
A wrong date in a sitemap is worse than no date at all.
If your system cannot produce accurate dates, leave lastmod out. That is allowed, and it is the honest option.
Where it goes, and what to do when it grows
A sitemap usually lives at the root of the site, for example /sitemap.xml. Most site builders create one automatically, so first check whether you already have one, and look at what is in it.
There are two ways to tell Google where it is.
1. A line in robots.txt. Add the full address of the sitemap to your robots.txt file. Any search engine that reads the file will find it.
Sitemap: https://www.example.com/sitemap.xml
2. Google Search Console. Open the Sitemaps section, enter the address and submit it. This has a bonus: Search Console tells you whether the file could be read and how many addresses it found.
I do both. They do not conflict.
Size limits and sitemap index files
One sitemap file may list up to 50,000 addresses and be up to 50 MB uncompressed. Most small sites are far below that. If you pass either limit, split the list into several files and name them in a sitemap index file:
<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<sitemap>
<loc>https://www.example.com/sitemap-pages.xml</loc>
</sitemap>
<sitemap>
<loc>https://www.example.com/sitemap-posts.xml</loc>
</sitemap>
</sitemapindex>
Then submit only the index file. Splitting can help below the limit too. With one file for pages, one for posts and one for products, it is easier to see which part of the site has a problem. On very large sites, a clean sitemap also helps Google spend its crawl budget on pages that matter.
One last point: a sitemap does not replace links. A page that appears in the sitemap but has no links pointing to it is still hard to reach. Pair this lesson with an internal linking plan you can reuse, and see the rest of the Technical SEO series for more.
Check your own site
Open your sitemap in a browser and look through it:
- Does every address load normally, with no redirects or errors?
- Is every listed page one you want to appear in search?
- Are there duplicates, filtered addresses or search result pages in the list?
- Are the lastmod dates true, or is every page stamped with the same date?
- Is the sitemap named in robots.txt and submitted in Search Console?
To check a single page from the list in more detail, run the free audit.