Back to guides

XML Sitemaps: Giving Crawlers a Clean Map of Your Site

A sitemap is a discovery aid rather than a ranking lever. How to keep yours discoverable and honest, and which URLs should never appear in it.

A sitemap is a discovery aid. It tells search engines which URLs you consider worth crawling and roughly when each last changed. On a small site with clean internal linking it changes little. On a large site, a new site, or one where useful pages sit 4 clicks from the homepage, it is often the difference between a page being found this week and next quarter. Two properties make a sitemap useful: it is discoverable, and it is honest. Discoverable means it is referenced from robots.txt or submitted directly, and it lives at a stable URL. Honest means every URL in it is one you would want indexed, which is to say canonical, returning 200, and open to crawling. The most common defect is a sitemap generated once and never regenerated, listing URLs that now redirect or 404 while omitting everything published since. The second is a sitemap containing every URL the CMS can produce, including tag archives, paginated duplicates and parameterized variants. Both teach search engines to trust the file less. A healthy sitemap can be found, parses as valid XML, and uses index and child files correctly. Every URL it lists should also agree with your own directives, because a file that nominates blocked, noindexed or non-canonical URLs is arguing with itself, and search engines settle that argument by trusting the file less.

SEOReport's paid diagnosis reviews this across the pages of your own site, shows the evidence behind every finding, and ranks the fixes by priority. See plans and pricing.

Get the complete diagnosis of your site

An evidence-backed report and a prioritized action plan, on a plan with monthly credits.