Entrovix AI

Sitemap Generator

Crawls your site and builds sitemap.xml — leaving out noindex pages and non-canonical duplicates, because a sitemap is a list of pages you want indexed.

Crawl a site

Follows internal links from the address you give and builds sitemap.xml from what it finds.

A sitemap is a list of pages you want indexed. Including a noindex page, or a duplicate whose canonical points elsewhere, contradicts the page itself — and Search Console reports it back to you as an error. Most free generators list everything they can reach.

Why use it

Built to be genuinely useful

Excludes what should not be there

noindex pages and pages whose canonical points elsewhere are left out, with the reason listed.

Obeys robots.txt

Reads your robots.txt and skips what it blocks, so the sitemap does not contradict it.

Every exclusion explained

A list of what was left out and why — so you can tell a deliberate decision from a bug.

Ready to submit

Valid XML, correctly escaped, downloadable as a file.

How it works

Three steps

  1. 1

    Enter your home page URL.

  2. 2

    Crawl — internal links are followed automatically.

  3. 3

    Download sitemap.xml and submit it in Search Console.

What a sitemap is for

A sitemap is a list of the pages you want in search results, in a format Google reads directly. It does not guarantee indexing and it is not a ranking factor — it is a way of making sure nothing is missed.

It matters most for pages that internal links reach poorly: a new site with few links, a large site where deep pages sit five clicks from the home page, or pages that were only ever linked from a menu a crawler renders inconsistently.

For a small, well-linked site it changes very little, and that is worth saying plainly. Google finds your pages through links regardless. The sitemap is insurance, and cheap insurance, but not a lever.

Why exclusions are the actual feature

Most free generators list every URL they can reach. That produces a file containing your noindex pages, your tag archives, your paginated duplicates and every URL whose canonical points somewhere else — and each of those is a contradiction: the sitemap says index this, the page says do not.

Search Console reports those back to you as errors, under headings like "Submitted URL marked noindex" and "Alternate page with proper canonical tag", and the usual response is confusion rather than a fix.

This crawler reads each page before including it. A page marked noindex is left out. A page whose canonical points elsewhere is left out. Anything robots.txt blocks is left out. And every exclusion is listed with its reason, so you can see whether it was a decision or a mistake in your own markup.

The limits of a browser-based crawl

This stops at 60 pages. That covers most brochure sites, small businesses and young blogs completely, and it will not cover a shop with 900 products. The tool says when it has stopped rather than handing you a partial file that looks complete.

It reads the HTML your server returns and does not run JavaScript, so a site that renders its navigation entirely in the browser will yield fewer pages here than Google eventually finds. If your links only exist after a script runs, that is worth knowing for its own sake.

For a large site, use your CMS's sitemap — WordPress generates one automatically, Shopify and Wix both do — because it is built from the database rather than from crawling, and it updates itself when you publish. This is for sites that have no such thing.

FAQ

Questions people ask

Need a tool like this for your business?

We build internal tools, dashboards and automation that fit how your team works.

See Our Services