Sitemap Indexed vs. Submitted
-
My sitemap has been submitted to Google for well over 6 months and is updated frequently, a total of 979 URLs have been submitted by only 145 indexed. What can I do to get Google to index them all?
-
SF finding 'useless' links is actually part of its purpose, if you believe they're useless you should be asking why they're there. Your XML sitemap should have nothing but clean URLs; 200 response codes and not canonicalized to another URL. The problem isn't that you have category URLs, it's that those (like the one in my previous example) have a canonical tag that points elsewhere. Anytime this is the case, the URL is considered un-indexable. You can see the proof of this by doing a Google search for "https://www.interstellarstore.com/meteorite-jewelry/meteorite-necklaces", I just checked and this URL isn't in the index.
You mentioned the age in your original comment that your XML sitemap had been submitted for well over 6 months, that's where I got the age from, maybe I misunderstood?
You have no reason to not trust SF, it's one of the most valuable tools in an SEO's toolbox. I've used it for 5+ years to create hundreds of sitemaps and countless other SEO tasks with no problem in providing reliable, accurate data points.
-
Hi Logan,
I tried using Screaming Frog but it kept finding useless links, so I wrote the sitemap myself and I update it manually, I updated it only this morning. What makes you think it is over 6 months since an update?
I was told on Moz in an earlier post that having all of the category links, not just the canonical ones, wasn't a problem, is this not the case?
Every link in the sitemap should work fine, I wrote it by copy and pasting the links directly from my site. I have no trust in Screaming Frog.
-
Hi,
I poked around a bit on your sitemap and noticed a couple things:
- You've got URLs on there that have canonicals to another page. For example:This page https://www.interstellarstore.com/meteorite-jewelry/meteorite-necklaces has a canonical tag that points here https://www.interstellarstore.com/meteorite-necklaces.
- A bunch of the URLs in your sitemap redirect elsewhere or have no response - I got 13% through crawling your XML sitemap with Screaming Frog and there were zero 200 response code URLs, not good.
Both of these things combined are causing a discrepancy in the amount of submitted URLs vs. indexed URLs. If you use Screaming Frog to create your XML sitemap it's quite easy to have only clean URLs in there. You can easily remove all URLs that are not 200 status and by default Screaming Frog will exclude any URL that canonicalizes to another URL.
Also, as a side note, you should be updating your XML sitemap more frequently, a 6 month old sitemap for an ecommerce site is far too old with new products being added and products dropping off.
Got a burning SEO question?
Subscribe to Moz Pro to gain full access to Q&A, answer questions, and ask your own.
Browse Questions
Explore more categories
-
Moz Tools
Chat with the community about the Moz tools.
-
SEO Tactics
Discuss the SEO process with fellow marketers
-
Community
Discuss industry events, jobs, and news!
-
Digital Marketing
Chat about tactics outside of SEO
-
Research & Trends
Dive into research and trends in the search industry.
-
Support
Connect on product support and feature requests.
Related Questions
-
Any way to force a URL out of Google index?
As far as I know, there is no way to truly FORCE a URL to be removed from Google's index. We have a page that is being stubborn. Even after it was 301 redirected to an internal secure page months ago and a noindex tag was placed on it in the backend, it still remains in the Google index. I also submitted a request through the remove outdated content tool https://www.google.com/webmasters/tools/removals and it said the content has been removed. My understanding though is that this only updates the cache to be consistent with the current index. So if it's still in the index, this will not remove it. Just asking for confirmation - is there truly any way to force a URL out of the index? Or to even suggest more strongly that it be removed? It's the first listing in this search https://www.google.com/search?q=hcahranswers&rlz=1C1GGRV_enUS753US755&oq=hcahr&aqs=chrome.0.69i59j69i57j69i60j0l3.1700j0j8&sourceid=chrome&ie=UTF-8
Intermediate & Advanced SEO | | MJTrevens0 -
Best Sitemap for Large Website
i have more than 3500 pages on my website. Please let me know the best sitemap plugin for my website.
Intermediate & Advanced SEO | | Michael.Leonard1 -
Blog - subdomain vs. subfolderq
Hi everyone I work on an ecommerce site and I'm trying to get more content together for the site & blog. The development team want to put the blog we have on a subdomain of our site, my question is - what is better for SEO Subfolder vs. subdomain I've read a couple of articles to say subfolder is better and a subdomain needs a lot of management to build up authority itself? Thanks!
Intermediate & Advanced SEO | | BeckyKey0 -
Best practice to prevent pages from being indexed?
Generally speaking, is it better to use robots.txt or rel=noindex to prevent duplicate pages from being indexed?
Intermediate & Advanced SEO | | TheaterMania0 -
Reducing Booking Engine Indexation
Hi Mozzers, I am working on a site with a very useful room booking engine. Helpful as it may be, all the variations (2 bedrooms, 3 bedrooms, room with a view, etc, etc,) are indexed by Google. Section 13 on Search Pagination in Dr. Pete's great post on Panda http://www.seomoz.org/blog/duplicate-content-in-a-post-panda-world speaks to our issue, but I was wondering since 2 (!) years have gone by, if there are any additional solutions y'all might recommend. We want to cut down on the duplicate titles and content and get the useful but not useful for SERPs online booking pages out of the index. Any thoughts? Thanks for your help.
Intermediate & Advanced SEO | | Leverage_Marketing0 -
No index.no follow certain pages
Hi, I want to stop Google et al from finding a some pages within my website. the url is www.mywebsite.com/call_backrequest.php?rid=14 As these pages are creating a lot of duplicate content issues. Would the easiest solution be to place a 'Nofollow/Noindex' META tag in page www.mywebsite.com/call_backrequest.php many thanks in advance
Intermediate & Advanced SEO | | wood1e19680 -
Why do my https pages index while noindexed?
I have some tag pages on one of my sites that I meta noindexed. This worked for the http version, which they are canonical'd to but now the https:// version is indexing. The https version is both noindexed and has a canonical to the http version, but they still show up! I even have wordpress set up to redirect all https: to http! For some reason these pages are STILL showing in the SERPS though. Any experience or advice would be greatly appreciated. Example page: https://www.michaelpadway.com/tag/insurance-coverage/ Thanks all!
Intermediate & Advanced SEO | | MarloSchneider0 -
Best Product URL For Indexing
My proposed URL: mydomain.com/products/category/subcategory/product detail Puts my products 4 levels deep. Is this too deep to get my products indexed?
Intermediate & Advanced SEO | | waynekolenchuk0