How to Build an XML Sitemap: Indexable Pages, Last-Modified Dates, and Duplicate URL Rules
Sitemaps often have one of two extremes: none at all, or every parameter page, redirect, 404, noindex page, and duplicate language URL. The latter continuously sends contradictory search signals.
The principle is simple: list only pages you want indexed with that URL as the canonical version.
01 Decide Which Pages Should Be Indexed
A corporate website commonly includes home, service, product, solution, case-study, article, and necessary brand pages. Login, search results, filter parameters, test, Thank You, and admin pages generally do not belong.
Sitemap inclusion must align with robots, noindex, canonical, and page status. A URL listed while marked noindex complicates diagnosis.

Sitemap URL Check
| URL Status | Recommended to Include? | Treatment |
|---|---|---|
| 200, Indexable Canonical Page | Yes | Use a complete HTTPS absolute URL |
| 301 or 302 Redirect | No | Replace with the final target URL |
| 404/410 | No | Remove from the Sitemap and check internal links |
| Noindex Page | No | Keep accessible but do not list |
| Canonical Points Elsewhere | Usually not | List only the canonical version |
| Paginated or Filtered Page | Depends on value | Include only with independent indexable value |
02 Use Absolute, Consistent, Canonical URLs
URLs need protocols and complete domains with consistent www, HTTPS, trailing-slash, and case rules. A Sitemap does not repair URL disorder; normalize routing before generation.
Multilingual pages must align with hreflang and canonical, avoiding duplicate language URLs across Sitemaps.

03 Update lastmod Only for Substantive Changes
Modification dates can inform crawling only when credible. Setting every page to today on every build makes the field meaningless.
Update after substantive changes to body content, product data, structured data, or critical media; footer-year and minor style changes generally do not qualify.
04 Split Large Sites by Type for Easier Monitoring
A Sitemap has URL and file-size limits and needs an index file when exceeded. Even smaller sites can split products, articles, case studies, and languages to simplify Search Console monitoring.
Split files to clarify ownership and diagnose categories, not merely to create more files.

05 Run Quality Checks Even After Automatic Generation
CMSs and frameworks know routes but not necessarily business indexing intent. Test pages, old content, and error states may still appear.
Sample before launch and scan status codes, canonicals, noindex, and duplicates regularly. Verify new sections enter the correct file.
06 Compare Discovery and Indexing After Submission
In Search Console, compare discovered, crawled, and indexed counts, focusing on submitted-but-not-indexed pages, redirect errors, and duplicate canonicalization.
Do not resubmit repeatedly after one indexing miss. Check distinct value, internal-link support, and server reliability first.
Frequently Asked Questions
How Many URLs Can One Sitemap Contain?
Google's limit is 50,000 URLs or 50 MB uncompressed per Sitemap. Split larger sets and use a Sitemap index.
Does Every Image Need an Image Sitemap?
Not necessarily. Standard sites can use crawlable images in pages; image-heavy sites or those relying on image search can consider an extension or separate Sitemap.
Are priority and changefreq Needed?
Google does not rely on them for crawl priority. Canonical URLs and credible lastmod values matter more.
Can robots.txt Include the Sitemap Address?
Yes, and submission in Search Console is also recommended. A Sitemap line in robots.txt enables discovery; it is not an access rule.
Can a Page Outside the Sitemap Still Be Indexed?
Yes. Search engines may find it through internal or external links, but important indexable pages should be listed for discovery and monitoring.
| Service | View |
|---|---|
| Related Services | View Service Details |
| Project Inquiry | Contact JVDS Design Studio |
| Design and Website Development Articles | View Service Details |