A machine-readable list of your important URLs, submitted to help discovery.
An XML sitemap is a structured file listing the URLs you want search engines to know about, optionally with last-modified dates. It helps crawlers discover pages, especially on large sites, deep pages, or sites with weak internal linking. It is a discovery aid, not a guarantee of crawling or indexing.
A sitemap should contain only canonical, indexable URLs that return 200. Filling it with redirects, noindex pages, or non-canonical URLs sends mixed signals and can slow discovery of the pages that matter.
Sitemaps are most valuable for large or poorly linked sites, and for surfacing fresh content quickly via last-modified dates. Submitting one in Search Console also gives you an indexation report to spot excluded pages.
Keep it clean and current, and let it reflect the true canonical structure of the site.
An XML sitemap is how you hand Google a clean, prioritised list of the URLs you actually want indexed, along with signals like last-modified dates. It does not guarantee indexing and it is not a substitute for good internal linking, but on large or poorly-linked sites it is a genuine discovery aid — it surfaces pages that are buried deep in the architecture or newly published. Its most useful modern role is diagnostic: submitting a sitemap in Search Console lets you compare "submitted" vs "indexed" and see exactly how much of your intended set Google is actually keeping.
A sitemap should contain only canonical, indexable, 200-status URLs — the pages you genuinely want in the index. A common mistake is dumping every URL in, including redirects, noindexed pages, non-canonical duplicates and 404s, which sends Google mixed signals and dilutes the sitemap's value. For large sites, split into multiple sitemaps under a sitemap index (the 50,000-URL / 50MB limits per file), and keep `lastmod` dates honest so Google can prioritise genuinely-changed pages.
A 40,000-page marketplace wants to know why only 26,000 pages are indexed. Submitting a clean XML sitemap turns the question into a diagnostic: Search Console's sitemap report shows 40,000 submitted vs 26,000 indexed, and the Page Indexing report attributes most of the gap to 'Crawled - currently not indexed' on thin, near-duplicate listing pages. The sitemap did not fix the gap, but it made it measurable and located it. The response is to consolidate or improve the thin listings and remove non-canonical URLs from the sitemap so it reflects only pages worth indexing. Kept accurate over time — canonical, 200-status URLs only, honest lastmod dates — the sitemap becomes an ongoing monitor of how much of your intended index Google is actually keeping, which is arguably its most valuable modern role.
Part of our defined terms knowledge graph — browse every entry in this branch.
How many pages a search engine will crawl on your site, and how often.
Whether a page is stored in a search engine’s index and eligible to rank.
A root-level file that tells crawlers which URLs they may or may not fetch.
A controlled experiment comparing two versions to see which performs better.
The visible, clickable words in a hyperlink, used as a relevance signal.
Common questions
Straight answers on how this fits your marketing and build.
Still have questions? Talk to a specialist