YouTube Video
Sitemaps have been a core element of web development and search engine optimization for two decades, yet Search Console warnings like "Couldn't fetch" continue to generate concern among site owners. Are XML sitemaps still necessary with modern search crawlers, or can search engines and AI bots figure out site structure on their own? In this episode of Search off the Record, Google Search Relations team members Martin Splitt and John Mueller break down the history, technical specifications, and some misconceptions surrounding sitemaps, RSS feeds, LLMs.txt, canonicalization signals, and Search Console indexing behaviors.
In this episode, you’ll learn: The Origins of Sitemaps: How early web crawling challenges led to the sitemap standard and how building a sitemap generator started John Mueller's career at Google.
Obsolete vs. Active XML Fields: Why Google dropped support for the priority and changefreq tags, and how the lastmod date is actually evaluated.
XML Sitemaps vs. RSS Feeds: How RSS feeds function as lightweight sitemaps for recent updates and why AI training crawlers utilize both.
Technical Limits & Index Files: The exact URL (50,000) and file size (50MB uncompressed) limits, alongside best practices for sitemap index files and robots.txt declarations.
HTML Sitemaps & LLMs.txt: Why user-facing HTML maps and Markdown-based LLMs.txt files cannot replace structured XML sitemaps for search engines.
Troubleshooting "Couldn't Fetch": How host load management and perceived site quality (crawl demand) trigger non-technical fetching errors in Search Console.
Key takeaways for SEOs & developers: Keep Sitemaps Enabled by Default: Even for small websites or dynamic CMS setups, keeping auto-generated sitemaps active carries no downside and assists with discovery as sites grow.
Provide Accurate lastmod Timestamps: Avoid setting all page timestamps to the current date; Google evaluates date reliability and will ignore lastmod signals if they are inaccurate or abused.
Submit Canonical URLs Only: Always place clean, canonical versions of URLs inside your sitemap file rather than parameter-tagged tracking URLs to help guide Google’s canonical selection.
"Couldn't Fetch" Isn't Always a Syntax Error: If Search Console reports "Couldn't fetch" on a valid, accessible XML file, the cause is often host load throttling or low crawl demand linked to perceived site quality.
Chapters 00:00 – Introduction & Welcome
00:57 – Sitemaps History & John Mueller's Path to Google
03:42 – Why Search Engines Started Using XML Sitemaps
04:46 – Deprecated Fields: Priority, Change Frequency, and lastmod
06:36 – Sitemaps for Small Sites vs. High-Volume News Outlets
08:50 – Canonical Selection & Parameterized URLs in Sitemaps
11:13 – Comparing RSS Feeds and XML Sitemaps
12:58 – Size Limits, Compression, Index Files, and robots.txt
14:15 – Custom File Naming, Privacy, and AI Crawlers
16:54 – Why HTML Sitemaps Don't Replace XML Sitemaps
19:06 – Will LLMs.txt Replace Sitemaps for AI?
22:05 – Decoding Search Console's "Couldn't Fetch" Error
24:16 – Hreflang, Image, and Video Sitemap Extensions
25:14 – Key Takeaways & Outro
Resources Mentioned:
Official Google Search Central Documentation: https://developers.google.com/search
Google Search Console: https://search.google.com/search-console
Don't forget to like and subscribe to the podcast on your favorite platform to catch every behind-the-scenes episode from the Search Relations team!
Episode transcript → https://goo.gle/sotr114-transcript
Listen to more Search Off the Record → https://goo.gle/sotr-yt
Subscribe to Google Search Channel → https://goo.gle/SearchCentral
Search Off the Record is a podcast series that takes you behind the scenes of Google Search with the Search Relations team.
#SOTRpodcast #SEO #GoogleSearch #SearchConsol#SEO #WebDevelopment #GoogleSearchConsole #Sitemaps #SearchEngineOptimization #GoogleSearchRelations #LLMstxt #TechnicalSEO #SearchOffTheRecord
Speakers: Martin Splitt, John Mueller
In this episode, you’ll learn: The Origins of Sitemaps: How early web crawling challenges led to the sitemap standard and how building a sitemap generator started John Mueller's career at Google.
Obsolete vs. Active XML Fields: Why Google dropped support for the priority and changefreq tags, and how the lastmod date is actually evaluated.
XML Sitemaps vs. RSS Feeds: How RSS feeds function as lightweight sitemaps for recent updates and why AI training crawlers utilize both.
Technical Limits & Index Files: The exact URL (50,000) and file size (50MB uncompressed) limits, alongside best practices for sitemap index files and robots.txt declarations.
HTML Sitemaps & LLMs.txt: Why user-facing HTML maps and Markdown-based LLMs.txt files cannot replace structured XML sitemaps for search engines.
Troubleshooting "Couldn't Fetch": How host load management and perceived site quality (crawl demand) trigger non-technical fetching errors in Search Console.
Key takeaways for SEOs & developers: Keep Sitemaps Enabled by Default: Even for small websites or dynamic CMS setups, keeping auto-generated sitemaps active carries no downside and assists with discovery as sites grow.
Provide Accurate lastmod Timestamps: Avoid setting all page timestamps to the current date; Google evaluates date reliability and will ignore lastmod signals if they are inaccurate or abused.
Submit Canonical URLs Only: Always place clean, canonical versions of URLs inside your sitemap file rather than parameter-tagged tracking URLs to help guide Google’s canonical selection.
"Couldn't Fetch" Isn't Always a Syntax Error: If Search Console reports "Couldn't fetch" on a valid, accessible XML file, the cause is often host load throttling or low crawl demand linked to perceived site quality.
Chapters 00:00 – Introduction & Welcome
00:57 – Sitemaps History & John Mueller's Path to Google
03:42 – Why Search Engines Started Using XML Sitemaps
04:46 – Deprecated Fields: Priority, Change Frequency, and lastmod
06:36 – Sitemaps for Small Sites vs. High-Volume News Outlets
08:50 – Canonical Selection & Parameterized URLs in Sitemaps
11:13 – Comparing RSS Feeds and XML Sitemaps
12:58 – Size Limits, Compression, Index Files, and robots.txt
14:15 – Custom File Naming, Privacy, and AI Crawlers
16:54 – Why HTML Sitemaps Don't Replace XML Sitemaps
19:06 – Will LLMs.txt Replace Sitemaps for AI?
22:05 – Decoding Search Console's "Couldn't Fetch" Error
24:16 – Hreflang, Image, and Video Sitemap Extensions
25:14 – Key Takeaways & Outro
Resources Mentioned:
Official Google Search Central Documentation: https://developers.google.com/search
Google Search Console: https://search.google.com/search-console
Don't forget to like and subscribe to the podcast on your favorite platform to catch every behind-the-scenes episode from the Search Relations team!
Episode transcript → https://goo.gle/sotr114-transcript
Listen to more Search Off the Record → https://goo.gle/sotr-yt
Subscribe to Google Search Channel → https://goo.gle/SearchCentral
Search Off the Record is a podcast series that takes you behind the scenes of Google Search with the Search Relations team.
#SOTRpodcast #SEO #GoogleSearch #SearchConsol#SEO #WebDevelopment #GoogleSearchConsole #Sitemaps #SearchEngineOptimization #GoogleSearchRelations #LLMstxt #TechnicalSEO #SearchOffTheRecord
Speakers: Martin Splitt, John Mueller