Documentation
Sitemap Scraper
Collect URLs from a website's sitemaps and download them.
Sitemap Scraper, the PirateSERP sitemap reader, collects every URL a public website lists in its XML sitemaps, including sitemap index files and gzipped sitemaps. Sitemap Scraper needs no API key.
Collect URLs
- Review Domain or URL, or enter a domain such as
example.com. The active project's business website can fill this field. - Click Scrape Sitemap.
- Review URLs listed and Sitemaps read, and check the result heading for a truncation note.
Sitemap Scraper reads the sitemaps that the site's robots.txt file declares, or /sitemap.xml when robots.txt declares none. The table shows URL and, when the sitemap supplies them, Last modified, Change frequency, and Priority. The table displays the first 500 URLs, and every download includes all collected URLs.
Download results
- Download CSV saves each URL with its available sitemap fields for a spreadsheet.
- Download TXT saves one URL per line.
- Download XML saves one combined sitemap that contains every collected entry.
A sitemap entry does not prove that a page is reachable or indexed. Use Site Auditor to crawl and inspect the pages.
If results are missing
"No sitemap was found at /sitemap.xml or in robots.txt" means the site publishes no sitemap at either location. "The sitemap could not be fetched" means the request failed, often because the site blocks automated requests. Check the domain and open its sitemap in your browser.
A result heading that ends with "truncated" means the site lists more URLs than one scrape collects; running the same scrape again returns the same set.
Sitemap Scraper keeps no scrape history, so download the results you need before you leave the page.