An XML sitemap is a list of the URLs you want crawlers to know about. lastmod is the date on each URL for the last significant change. Google’s guide to building a sitemap treats that list as a way to tell Google which pages you consider canonical and worth crawling. It is not a ranking file. A generator that stamps today’s date on every URL, and sets priority to 1.0 on all of them, is doing the two things Google has said not to bother with, or not to fake.
List the canonical URL that returns 200, and date it honestly. Leave out redirects, noindex pages, and tracking-parameter copies. Set lastmod when the article, its structured data, or its links changed. Do not set priority. Do not rebuild the file every night with a new date and the same pages.
What an XML sitemap is
The file is XML, UTF-8, with a urlset and one url block per address. The required child is loc, the absolute URL. lastmod is optional and, when it is true, useful. A small blog can live in one file. You put the file on the site, usually at /sitemap.xml, and you point to it from robots.txt with a Sitemap: line. You can also submit it in Search Console. Submitting it does not index the URLs. It tells Google where the list is.
Google will try to crawl the URLs as written. A relative path is the wrong value. https://northwind.example/async-standups is a location. /async-standups is not, inside a sitemap. Use the host you chose, HTTPS, and the same form your canonical tag uses. If the site lives on www, the sitemap should not list the bare host, and the reverse. The canonical tag page is how you pick that one address. The sitemap should repeat the choice, not offer a second one.
Order does not matter. Google says the order of URLs in the file is irrelevant. Alphabetical, newest first, or the order the CMS wrote them are all fine. What matters is that each loc is a page you would send a person to.
A sitemap is a hint about discovery and about which copy you prefer. Internal links are how a person moves, and they are also a crawl path. A URL that appears only in the sitemap can still be found by Google. A person on your site cannot find it if nothing links to it. Publish the article, link it from the page that mentions the topic, and list the canonical URL. The topic cluster guide is where those links belong in the sentence. The sitemap is the inventory.
Which URLs belong
Include a URL when it is the version you want crawled and, if it deserves to be a result, indexed. Leave out everything else.
| URL | In the sitemap? | Why |
|---|---|---|
| The article, HTTPS, canonical, 200 | Yes | This is the page. |
The same article with utm_ parameters |
No | It is a copy. Listing it fights the canonical. |
| An old slug that redirects | No | The redirect is the signal. The sitemap should not nominate the old address. |
| A thank-you page or a preview with noindex | No | The robots meta tag says it should not be a result. The sitemap should not argue. |
| A URL disallowed in robots.txt | No, if you still want it crawled | A disallow stops the fetch. Listing it does not unlock the crawl, and a noindex on a blocked URL is never read. |
| A tag archive that only repeats titles | Only if it is a real page you want found | A list with no job of its own is not a second article. |
| A 404 or a soft 404 | No | Do not ask Google to crawl a URL that does not resolve to the article. A 200 “not found” page is still the wrong code. |
After a merge or a slug change, remove the old path in the same edit as the redirect. Leaving it in the file is how both URLs stay nominated. The content audit decides which URL survives. The sitemap should match that decision the day you ship it, not at the next quarterly cleanup.
Home page and a few real category pages can be listed when they are pages with their own text. A category that is only a grid of posts is optional. If your CMS cannot tell when that grid last changed in a meaningful way, Google’s 2023 sitemap note says you can omit lastmod for that URL rather than invent one. Omitting the date is more honest than a date that moves whenever any post on the site is saved. Page 2 of a blog index is the same kind of list. Google finds it from the next-page link, which is pagination, so you do not have to list every page number. Do not list filtered copies of that list.
What lastmod is allowed to mean
Google uses lastmod when it is consistently and verifiably accurate, for example by comparing it with the page. The value is the last significant modification. Google’s sitemap documentation and the June 2023 post on lastmod and the sitemap ping endpoint use the same line: a change to the main content, the structured data, or the links is generally significant. A change to the copyright date in the footer is not. A sidebar tweak is not.
If the file says every article changed today, and the articles did not, Google can stop trusting the field. That is the failure of a generator that writes lastmod as the time the sitemap was rendered. The render is not an edit. Store the date of the edit, and write that date into the sitemap when you publish the edit.
The format is W3C datetime. 2026-10-06 is valid. A full timestamp with a timezone is better when you have a real time, because a site that changes often gives crawlers a clearer signal. Do not invent a time. Do not use a format Search Console will reject. After you submit the sitemap, the report is where a bad date format shows up.
The same date should agree with the page. If the article shows “Updated 6 October 2026,” dateModified in the article schema and lastmod in the sitemap should be that day, and only because the passage changed. The content refresh guide is what counts as that change. A sitemap date that moves alone is the date-only pattern, moved from the byline into the XML.
You do not have to put lastmod on every URL. Google says you can include it only where you are confident. A homepage that merely aggregates other pages is a common place to leave it off. An article you just rewrote is a place to include it.
Priority and changefreq
The sitemap protocol still defines priority and changefreq. Google’s June 2023 post says Google does not use either element at all. changefreq overlaps lastmod and was widely abused. priority is a subjective number. Sites set everything to 1.0, which made the number meaningless. Google’s internal view, stated in that post, is that the values do not reflect real priority.
You can leave the tags out. If a plugin writes priority 0.5 and changefreq weekly on every URL, you do not need a project to delete them, and you also should not spend time tuning them. They are not a crawl budget setting. They are not a ranking setting. The work is the URL list and the dates.
sitemaps.org says priority does not change how your pages compare with other sites. Even the protocol does not claim a ranking effect. Google goes further and ignores the field. A guide that tells you to set the pillar to 1.0 and the cluster pages to 0.8 is describing a dial that is not connected.
Size, encoding, and where the file lives
One sitemap file is limited to 50,000 URLs or 50MB uncompressed, whichever you hit first. Past that, split into several sitemaps and list them in a sitemap index. A company blog will not hit the cap. A site that emits a URL for every filter combination will, and those URLs usually should not have been listed. The cap is not a goal. A smaller file of real pages is the goal.
The file must be UTF-8. Host it on the site. Google says a sitemap affects descendant URLs of the directory it lives in, unless you submit it in Search Console. A sitemap at the root can cover the site, which is why the root is the usual place. A sitemap buried in a folder you never submit may not apply to the rest of the host.
Gzip is allowed if the uncompressed file still respects the limits. For a blog, plain XML is easier to open and check. Open it after a publish. Search for the new slug. Search for a slug you redirected and confirm it is gone.
Name the file something stable, such as sitemap.xml, and refer to that name from robots.txt. Renaming it every month strands the old location. If you split by type later, a sitemap index at the old URL can point at the parts, so the robots.txt line does not have to change.
Google retired the unauthenticated sitemap ping endpoint. The 2023 post is the announcement that ping was going away, and that lastmod is the signal they would rather have. Do not build a cron that pings a URL Google no longer wants. Submit the sitemap in Search Console once, keep the robots.txt line, and update lastmod when a page changes. Asking for indexing of one URL is a different tool, the one in the indexing setup, and it still does not override noindex.
What a sitemap does not do
A sitemap does not rank a page. It does not increase crawl budget, and compressing the file does not either. It can carry hreflang alternates if you choose the sitemap method instead of HTML link tags. Those child entries do not replace the URL list, and they do not detect the language. It does not force crawling on a schedule. It does not remove a URL. Taking a URL out of the sitemap does not noindex it. If another page links to it, Google can still know it exists. Removal is noindex, a redirect, or a password, depending on whether the URL should disappear, move, or stay private.
A sitemap does not fix a blocked crawl. If robots.txt disallows the article, the sitemap entry is a request Google may not be able to fulfill. Allow the crawl, then list the URL. The robots guide is that split. Do not disallow /blog and then wonder why the sitemap “is not working.”
A sitemap is not llms.txt. Google’s Search Central changelog for 15 June 2026 says an llms.txt file is not needed for Google Search and will not positively or negatively affect visibility or rankings there. Other systems may use one. Publishing it is optional and separate. It does not replace the XML list of your articles, and it does not make a passage quotable. The passage has to be on the page. Answer engines quote that text. The sitemap only helps a crawler find the URL where the text lives.
Rebuilding the sitemap on every request, with lastmod set to now, is also a small waste. A static file written when content changes is enough. The file is XML. It does not need JavaScript, a font, or a client-side render. Keep it that way. A sitemap that is an HTML page styled to look like XML is not a sitemap.
Image sitemaps and news sitemaps
An image sitemap can list pictures for Google Images. It is optional. It does not write alt text, and it does not put the image next to the sentence it explains. The image SEO guide is that work. If the image is already in the HTML of a listed article, you do not need a second file to “submit the image.” Add an image sitemap when you have a library of images that would not otherwise be discovered, not as a ritual on every blog post.
A news sitemap is a different format for news publishers. Evergreen guides do not belong in it. Do not create one because a plugin offered a checkbox. News features have their own rules. A company blog that marks every how-to as news is the same mismatch as NewsArticle on a procedure.
Video sitemaps follow the same idea. If you do not publish video, do not emit an empty video sitemap. Empty optional files are noise in Search Console.
An XML sitemap example
Northwind’s article lives at https://northwind.example/async-standups. It was rewritten on 6 October 2026. The sitemap contains that absolute URL once, with lastmod of 2026-10-06. The canonical tag on the page points at the same URL. The visible updated date is 6 October 2026. Internal links use that path, without a campaign parameter.
The file does not contain /status-meetings, which redirects. It does not contain /thank-you. It does not contain /async-standups?utm_source=newsletter. It does not contain priority or changefreq. It does not list a draft preview. When the next article ships, that URL is added with its own lastmod. The standup URL’s lastmod stays 6 October until the standup page changes again.
They do not run a job that rewrites the whole sitemap at midnight and sets every lastmod to the job’s clock. Search Console shows the sitemap as readable. A spot check of three URLs finds the dates they remember editing. That is the maintenance. A priority column in a spreadsheet is not.
How this shows up when BloGoose reads a site
BloGoose uses your sitemap to see which URLs already exist, so a new draft does not repeat a page you have. That only works if the sitemap is the real inventory. If it lists redirects, parameter copies, and noindex utilities, the planner is looking at addresses that are not articles. If it omits a post that is live and linked, the planner may propose a second URL for a job you already published. Clean the list, then connect it.
When BloGoose publishes, the new canonical URL should appear in the sitemap your site serves, with a lastmod that matches the publication, and without a noindex on the article. The destination CMS is what writes that file for a WordPress site. After the first post, open /sitemap.xml and the page source. One URL, one canonical, no noindex, and a date you believe.
A later refresh updates lastmod because the article changed, not because the sitemap was regenerated. If your host rebuilds the file on a timer, configure it to keep the stored modification time of each post. The timer is not the edit.
Questions about XML sitemaps
What is an XML sitemap?
An XML sitemap is a file that lists the URLs you want search engines to know about, usually with a lastmod date. It is a discovery hint. It does not rank the URLs, and it does not force Google to index them.
What is lastmod?
lastmod is the date of the last significant change to that URL. Google uses it when it is consistently accurate, compared with the page. A significant change is the main content, the structured data, or the links. A copyright-year edit is not.
Should you set priority and changefreq?
No. Google has said it does not use priority or changefreq. Setting every URL to priority 1.0 does not make those pages more important. Spend the effort on an accurate lastmod and on listing the right URLs.
Which URLs should be left out?
Leave out redirects, error pages, URLs with a noindex tag, and duplicate addresses such as tracking parameters. List the canonical URL that returns 200. A sitemap that also lists the old slug works against the redirect.
Does a sitemap replace internal links?
No. Crawlers also find pages from links. A URL that is only in the sitemap and never linked can still be discovered, but people cannot. Link the article from the page that mentions it. The sitemap is the list, not the only path.
Does llms.txt replace a sitemap?
No. Google’s Search documentation says an llms.txt file is not needed for Google Search and does not affect visibility or rankings there. Other systems may read it. Your XML sitemap is still the list of URLs you want crawled.
Point BloGoose at a sitemap that lists the articles you actually have.
Start the 1-day trial