I have two sitemaps which partly duplicate - one is blocked by robots.txt but can't figure out why!

McTaggart

Hi, I've just found two sitemaps - one of them is .php and represents part of the site structure on the website. The second is a .txt file which lists every page on the website. The .txt file is blocked via robots exclusion protocol (which doesn't appear to be very logical as it's the only full sitemap). Any ideas why a developer might have done that?

Chris.Menke

There are standards for the sitemaps .txt and .xml sitemaps, where there are no standards for html varieties. Neither guarantees the listed pages will be crawled, though. HTML has some advantage of potentially passing pagerank, where .txt and .xml varieties don't.

These days, xml sitemaps may be more common than .txt sitemaps but both perform the same function.

McTaggart

yes, sitemap.txt is blocked for some strange reason. I know SEOs do this sometimes for various reasons, but in this case it just doesn't make sense - not to me, anyway.

McTaggart

Thanks for the useful feedback Chris - much appreciated - Is it good practice to use both - I guess it's a good idea if onsite version only includes top-level pages? PS. Just checking nature of block!

Chris.Menke

Luke,

The .php one would have been created as a navigation tool to help users find what they're looking for faster, as well as to provide html links to search engine spiders to help them reach all pages on the site. On small sites, such sitemaps often include all pages of the site, on large ones, it might just be high level pages. The .txt file is non html and exists to provide search engines with a full list of urls on the site for the sole purpose of helping search engines index all the site's pages.

The robots.txt file can also be used to specify the location of the sitemap.txt file such as

sitemap: http://www.example.com/sitemap_location.txt

Are you sure the sitemap is being blocked by the robots.txt file or is the robots.txt file just listing the location of the sitemap.txt?

Welcome to the Q&A Forum

Browse the forum for helpful insights and fresh discussions about all things SEO.

I have two sitemaps which partly duplicate - one is blocked by robots.txt but can't figure out why!

Got a burning SEO question?

Browse Questions

Explore more categories

Related Questions

Can anyone please explain the real difference between backlinks, 301 links, and redirect links?which one is better to rank a website? i am looking for the help for one of my website

What do you add to your robots.txt on your ecommerce sites?

Default Robots.txt in WordPress - Should i change it??

How to solve outbound broken links? Those don't exist now?

Huge increase in server errors and robots.txt

Should I include www in url, or doesn't it matter?

New website won't rank for branded keywords in Google, but does in Bing

Should I disallow via robots.txt for my sub folder country TLD's?