Google’s Martin Splitt and John Mueller discussed the reasons why Search Console may display a “couldn’t fetch” error even though a website displays a valid XML sitemap. While Mueller acknowledged there sometimes may be a technical reason for that happening he also said that the actual reason is often a quality issue.
Google Acknowledges Search Console’s Inadequate Message
Splitt acknowledged that many people who receive the “couldn’t fetch” error have valid XML sitemaps that are properly linked from robots.txt and there are no technical reasons for why Google would be unable to fetch the sitemap. The takeaway here is that workers at Google are aware of the problem and the user frustration (more about this later).
Martin Splitt observed:
“Why does Search Console sometimes show “couldn’t fetch”, even though the sitemap XML is valid?
Because you can validate that. That’s the nice thing about such a structured format. So it’s valid, it’s publicly accessible, and it’s linked from the robots.txt. And yet Search Console says can’t fetch sometimes.”
Reason 1: Host Load Issues
Google’s John Mueller shared that they often see questions about the can’t fetch error message in forums and that there are two reasons for this error message.
The first reason is that sometimes Google really cannot access the sitemap because of server host load issues, which means that the server has too many incoming requests for pages and is unable to serve the requested resource. But it can also mean that the server can’t handle the crawling load.
Mueller explains:
“We’ve talked about that in the past, like how much Google systems are able to crawl from a website. And it could be the case that we don’t have any time to crawl this sitemap file because we’re too busy with other things. That can happen. Then we would also position that as “couldn’t fetch” because we didn’t have time to actually fetch it.”
Though Mueller didn’t mention this, the host load issue can also happen late at night when legit and non-legit crawlers hammer a website with thousands of requests all at once, at which point the server will give up and throw a 500 error response. The 500 server error response can be confirmed with Google’s Search Console where it lists 500 error responses and also in a server log file, if you have access to that.
Reason 2: Site Quality Issues
The second reason Mueller shared is what he called crawl demand but is really about Google perceiving that they don’t really need the content and deciding to skip it. He explained that this is a content quality issue. Mueller confusingly says it’s related to host load but I’m not sure I agree with him, judge for yourself.
Mueller shared:
“The other, also related to host load, is kind of the crawl demand side, where if our systems say, we don’t actually have any need to crawl a lot from this website, we’re just going to skip the sitemap file because we got enough already.
And the crawl demand is very often based on the perceived quality of a website. And that can have a really large impact on how much we crawl and index from a website. So it’s not purely a technical thing. Sometimes it’s that our systems assume that the overall quality of this website is not fantastic. Therefore, we’re not going to spend a lot of time crawling and indexing the content. Therefore, we’re not going to bother with the sitemap file at the moment.”
Reason 3: Maybe The Sitemap Is Not Needed
The third reason he shared is that sometimes Google doesn’t really need the sitemap.
Mueller explained:
“And if we see over time that the quality of the website improves significantly, then yes, we will go off and use that sitemap file, but maybe we just don’t want to. So it is not just a technical thing from my perspective or my site’s perspective.
It’s also, do we actually need the sitemap file and do we actually want the sitemap file as well?”
Takeaways
- Search Console’s “couldn’t fetch” sitemap error can be misleading.
A sitemap can be valid, publicly accessible, and properly linked while Search Console still reports that Google couldn’t fetch it. - Server load can prevent Google from fetching a sitemap.
If Googlebot cannot access the sitemap because the server is overloaded or Google has exhausted the site’s available crawl capacity, Search Console may report “couldn’t fetch.” - The error could be due to crawl demand.
Google may deliberately skip a sitemap when its systems determine there is little reason to crawl more of the site. - Site quality can influence whether Google bothers with the sitemap.
Mueller said perceived website quality can have a large impact on how much Google crawls and indexes, including whether it fetches the sitemap at all.
Lastly, although neither Martin Splitt or Mueller don’t mention it, this discussion highlights a big problem with Search Console in that the message (couldn’t fetch) does not match the actual reason. The consequence is that search console users end up confused and frustrated. The especially frustrating part about this is that Google obviously knows that search console’s message is unhelpful but they’re essentially shrugging and not doing anything about it.
Featured Image by Shutterstock/Chuenmanuse