DNS, Security & Integrations
Extract/Scrape URLs from Sitemap to CSV
Pull every URL out of a sitemap.xml (yours or a competitor's, where permitted) into a clean, usable CSV.
The Problem
You need a full list of a site's URLs — for an audit, a migration, or competitive research — and manually copying them out of a sitemap isn't realistic once it has more than a handful of entries.
About this problem
A sitemap.xml is a structured list of a site's URLs, and large sites split it into a sitemap index that points to many child sitemaps, which may also be compressed. Reading all of them by hand is unrealistic, and a simple copy-paste misses nested files or the lastmod dates. A script can fetch each file, follow the index, parse the XML and de-duplicate the results.
People look for this when preparing a site migration, auditing content, planning redirects or doing research on a site's structure. The output is a simple CSV that opens in a spreadsheet and can be used as a base for the next step.
What's Included
- Parsing the sitemap (including sitemap index files with multiple nested sitemaps)
- Extracting every URL, with last-modified dates if present
- Exporting to a clean, ready-to-use CSV
- Handling genuinely large sitemaps without missing entries
What's NOT Included
- Scraping content that requires bypassing a login or paywall
- Any scraping of a site whose terms of service explicitly prohibit it
- Ongoing/scheduled re-scraping unless set up as a separate tool
How It Works
- I fetch the sitemap URL you provide and detect whether it is a single sitemap or an index of sitemaps.
- I follow every child sitemap, handling compressed (gzip) files where present.
- I parse each entry and extract the URL and the last-modified date when it exists.
- I remove duplicates and check the total count against the entries in the source files.
- I export the data to a clean CSV with a header row and consistent encoding.
- I deliver the CSV and note any sitemap entries that could not be fetched.
In practice: you buy the service, send over whatever access or details the job needs, I investigate and do the work, and you confirm it's resolved before we call it done.
Frequently Asked Questions
- How do I get all URLs from a sitemap into Excel?
- By parsing the sitemap, including any nested ones, and exporting the URLs to CSV, which opens in Excel or Google Sheets. That is what this service does.
- Does it work with sitemap index files?
- Yes. It follows each nested sitemap and merges the results.
- Can you do it for a competitor's site?
- Sitemaps are public, but I only do it where the site's terms allow it.
- Will it bypass a login or paywall?
- No. Content behind a login or paywall is not covered.
- Can the extraction run on a schedule?
- Not as part of this service. Regular re-scraping would be set up as a separate tool.
- What if the site has no sitemap?
- Then there is nothing to parse. A crawler would be a different job.