Start with the pages you actually want people to reach
A sitemap is a list of URLs you want search engines to discover. It is not a list of every route your application happens to have. Before generating one, decide which public pages are useful to visitors and make sure each one returns the expected page rather than an error, redirect loop, or placeholder.
1. Keep the jobs separate
Sitemap
Helps search engines discover the public URLs you list. It does not guarantee that a URL will be crawled, indexed, or ranked.
robots.txt
Communicates crawler rules. It is not access control and should never be used to protect a private file, page, or secret.
Page controls
Use your application, authentication, permissions, and response headers for access control. Review the page itself when you need an indexing directive.
2. Review the source URL list before turning it into XML
A spreadsheet export, CMS report, or hand-maintained list can include duplicates, staging URLs, malformed dates, multiple hosts, and old paths. Find those problems in the source. Do not rely on a generated file to make a bad list useful.
Start from a CSV list
Check URLs and date values row by row, then download XML containing only the entries that passed review.
Create an XML sitemap from CSVCheck an existing XML file
Review core sitemap structure, URL fields, dates, duplicates, limits, and host consistency without fetching the listed pages.
Validate XML sitemap- Use the canonical public URL, including the intended protocol and host.
- Remove duplicate locations and URLs that are not meant for public discovery.
- Treat a malformed or missing date as a source-data issue, not a reason to invent a date.
3. Split large sitemaps deliberately
Large URL collections may need multiple sitemap files and one index file. Keep the original source list until you have confirmed the parts, their URL totals, and the links in the generated index. Do not manually cut XML at an arbitrary line because that can leave an incomplete entry behind.
Create upload-ready parts
Split one valid XML sitemap into URL-limited sitemap files and a sitemap index. The tool stops on malformed entries instead of silently dropping them.
Split an XML sitemapDo not use robots.txt as a security boundary
A disallow rule is a crawl instruction, not a lock. The address can still be known, linked, copied, or requested directly. Protect non-public content with real access controls, and remove it from public locations rather than relying on a crawler rule.
4. Write crawler rules as a reviewable policy
Every Allow or Disallow rule should have a clear reason. If you are unsure whether a rule blocks a useful public section, keep the change small and test it against the final URL paths. Include a sitemap location only after the destination file is publicly reachable.
Build and check crawler rules
Create groups, Allow and Disallow rules, and sitemap URLs with live validation. Read the guidance for each decision instead of applying a preset without review.
Create robots.txt5. Check the files after deployment
Your local file is not the final result. After deployment, request the public sitemap and robots.txt URLs in a normal browser or command-line client. Then inspect a few representative URLs from the sitemap, including the first, last, and a recently changed page.
- Confirm the public file returns successfully and is not replaced by an application error page.
- Confirm sitemap index links point to the intended public files.
- Check that robots.txt names the intended sitemap location.
- Use a search-engine webmaster tool for submission and indexing feedback; a valid file alone cannot promise a result.
When the files are ready
Publish the reviewed files, retain the source list, and record why a crawl rule exists. When the site changes, update only the parts affected by that change and repeat the public-file check.
Questions this checklist answers
Does submitting a sitemap guarantee indexing?
No. A sitemap helps discovery. Search engines decide whether and when to crawl, index, or show an eligible page.
Can I hide a private page with robots.txt?
No. Use authentication and real access controls for private content. A robots rule does not prevent direct access to a URL.
Should every URL in my application be in the sitemap?
No. List public, useful canonical pages that you want search engines to discover. Keep temporary, duplicate, error, and access-controlled URLs out of the source list.
