What would be considered a bad ratio to determine Index Bloat?
-
I am using Annie Cushing's most excellent site audit checklist from Google Docs. My question concerns Index Bloat because it is mentioned in her "Index" tab.
We have 6,595 indexed pages and only 4,226 of those pages have received 1 or more visits since January 1 2013.
Is this an acceptable ratio? If not, why not and what would be an acceptable ratio? I understand the basic concept that "dissipation of link juice and constrained crawl budget can have a significant impact on SEO traffic." [Thanks to Reid Bandremer http://www.lunametrics.com/blog/2013/04/08/fifteen-minute-seo-health-check/#sr=g&m=o&cp=or&ct=-tmc&st=(opu%20qspwjefe)&ts=1385081787]
If we make this an action item I'd like to have some idea how to prioritize it compared to other things that must be done. Thanks all!
-
Hi EGOL,
Wow, thank you so very much. This is one of the best answers I've ever received, probably the best, here in Q & A. Your thoughtful comments and suggestions are so appreciated. Honestly, you gave me a check list of things that have potential to be pure gold for us if we act on them.
Yes, you are correct, this is the site that had many issues with content being under tabs. It's also got a tremendous amount of duplicate and thin content issues, in addition to orphaned pages. Progress has been coming along, slowly and surely, but having your comments, and having them be so specific, pointed and concise are something I can take to my team and say "Here's an awesome check list of things that we can actually address right now, without re-platforming the site [you know, there are always people who think that the root of all a site's problems is the platform that it's on...pure mythology]."
I hope many others find your check list useful. Combined with Annie's audit spreadsheet in Google docs, I feel like I have the tools I need to go to battle and help this site fulfill its potential. Nearly every point you mentioned struck a chord. Better yet, now that I know my way around the "guts" of this homegrown CMS, I feel like I can actually make the necessary changes.
Egol, I really can't thank you enough.
-
I totally agree Keri. Every word Egol wrote , to me, is worth its weight in gold. I think this may be the best response I have ever received here in Q & A.
-
If only people realized how much good information members drop in Q&A...
Once again, thanks for this EGOL!
-
From my experience, that is a frightening number of pages that have not received a visit. I would definitely be taking some type of action. This hits to me like a site in very bad health. I have lots of little pages on a weak little site that get a lot more traffic than none since January. This would be high on my priority list of things to solve. Solving this could bring major income so this is potential opportunity as much as it is a problem.
To diagnose, I would check.... I know you and suspect that you have looked at all of these but just making a list, just in case.
A) Duplicate content problem? Does this site have lots of pages with very similar other pages on the same site. Does the company have another site that is running the same product descriptions? Does the site run product descriptions that are used from a datafeed supplied to vendors? Are affiliates using the same content? Have other websites stolen the content?
B) Have you been scraped and republished by a strong website? Just one is all it would take. A strong site was once scraping and republishing some of my short content pages and that killed the traffic into a section of my site. As soon as I asked them to stop traffic was back within days. One site can hurt you like that or numerous small sites - even minor sites in Asia can do this.
C) Lots of thin content? Do you have a lot of pages that might only have two or three unique sentences? Google could be disrespecting your entire site because of this.
D) Technical problem? I would be looking at robots.txt and .htaccess, noindex, badly coded links, content management system causing duplicated title tags or other problems? Faulty analyitics that make it look like these pages are not getting traffic when really they are.
E) Content cannibalization? Lots of separate pages for red widgets that are being filtered from the SERPs.
F) Inadequate linkjuice? This is not a huge site but not a small one. Does it have a nice amount of linkjuice coming in?
G) Does this site have pages that are really deeeeep down in the linkstructure? Many clicks down? Fix that either with a new linkstructure or some kickass powerful links that hit nodes deep in the site to force spiders down. I would solve with linkstructure.
H) This isn't the site that had all of the content behind tabs that I remember from a while ago? (My memory is really bad so it might not even be your site.) If you have pages like that I would get rid of those tabs immediately. I have a personal opinion that Google does not treat content hidden behind tabs as well as content that is out in the open.
I) Are there a lot of other sites - strong ones - publlishing very similar pages - like product description pages - competing for the same keywords. If that is the case you could be crowded out of the SERPs and receiving no traffic on these pages.
J) Does this site have a bad history? Does it have something that might be causing a penalty or filtering?
After doing all of that you might have something that is really worth fixing. If you can't identify the problem I would be slashing, hatcheting those pages from the site right away.
Got a burning SEO question?
Subscribe to Moz Pro to gain full access to Q&A, answer questions, and ask your own.
Browse Questions
Explore more categories
-
Moz Tools
Chat with the community about the Moz tools.
-
SEO Tactics
Discuss the SEO process with fellow marketers
-
Community
Discuss industry events, jobs, and news!
-
Digital Marketing
Chat about tactics outside of SEO
-
Research & Trends
Dive into research and trends in the search industry.
-
Support
Connect on product support and feature requests.
Related Questions
-
Why Are Some Pages On A New Domain Not Being Indexed?
Background: A company I am working with recently consolidated content from several existing domains into one new domain. Each of the old domains focused on a vertical and each had a number of product pages and a number of blog pages; these are now in directories on the new domain. For example, what was www.verticaldomainone.com/products/productname is now www.newdomain.com/verticalone/products/product name and the blog posts have moved from www.verticaldomaintwo.com/blog/blogpost to www.newdomain.com/verticaltwo/blog/blogpost. Many of those pages used to rank in the SERPs but they now do not. Investigation so far: Looking at Search Console's crawl stats most of the product pages and blog posts do not appear to be being indexed. This is confirmed by using the site: search modifier, which only returns a couple of products and a couple of blog posts in each vertical. Those pages are not the same as the pages with backlinks pointing directly at them. I've investigated the obvious points without success so far: There are a couple of issues with 301s that I am working with them to rectify but I have checked all pages on the old site and most redirects are in place and working There is currently no HTML or XML sitemap for the new site (this will be put in place soon) but I don't think this is an issue since a few products are being indexed and appearing in SERPs Search Console is returning no crawl errors, manual penalties, or anything else adverse Every product page is linked to from the /course page for the relevant vertical through a followed link. None of the pages have a noindex tag on them and the robots.txt allows all crawlers to access all pages One thing to note is that the site is build using react.js, so all content is within app.js. However this does not appear to affect pages higher up the navigation trees like the /vertical/products pages or the home page. So the question is: "Why might product and blog pages not be indexed on the new domain when they were previously and what can I do about it?"
Technical SEO | | BenjaminMorel0 -
AJAX and High Number Of URLS Indexed
I recently took over as the SEO for a large ecommerce site. Every Month or so our webmaster tools account is hit with a warning for a high number of URLS. In each message they send there is a sample of problematic URLS. 98% of each sample is not an actual URL on our site but is an AJAX request url that users are making. This is a server side request so the URL does not change when users make narrowing selections for items like size, color etc. Here is an example of what one of those looks like Tire?0-1.IBehaviorListener.0-border-border_body-VehicleFilter-VehicleSelectPanel-VehicleAttrsForm-Makes We have over 3 million indexed URLs according to Google because of this. We are not submitting these urls in our site maps, Google Bot is making lots of AJAX selections according to our server data. I have used the URL Handling Parameter Tool to target some of those parameters that are currently set to let Google decide and set it to "no urls" with those parameters to be indexed. I still need more time to see how effective that will be but it does seem to have slowed the number of URLs being indexed. Other notes: 1. Overall traffic to the site has been steady and even increasing. 2. Google bot crawls an average of 241000 urls each day according to our crawl stats. We are a large Ecommerce site that sells parts, accessories and apparel in the power sports industry. 3. We are using the Wicket frame work for our website. Thanks for your time.
Technical SEO | | RMATVMC0 -
Could schema.org and GoodRelations be bad for SEO?
One of my clients is going through a redesign and I am considering implementing schema.org and GoodRelations as it is an e-commerce website. The site sells cutting edge products and competes with some of the top tech blogs for rankings on the first page. Essentially, this means that e-commerce product listings are competing with news stories. It is becoming more and more difficult to rank as Google puts more emphasis on news over products in the serps, especially prior to a product release. My concern is that in implementing schema.org and GoodRelations, detailing to seach engines that this is in-fact a product page and not news could harm rankings. What opinions do others have on this?
Technical SEO | | pugh0 -
/index.php/ page
I was wondering if my system creates this page www my domain com/index.php/ is it better to block with robot.txt or just canonize?
Technical SEO | | ciznerguy0 -
Google indexing tags help
Hey everyone, So yesterday someone pointed out to me that Google is indexing tags and that will likely hurt search engine results. I just did a "site:thetechblock.com" and I notice that tags are still being pulled. http://d.pr/i/WmE6 Today, I went into my Yoast settings and checked "noindex,follow" tags in the Taxomomies settings. I just want to make sure what I'm doing is right. http://d.pr/i/zmbd Thanks guys
Technical SEO | | ttb0 -
Instant Indexing
I've been working on a site for a while now, methodically building content and building trust and authority. Lately I've noticed that anything I publish there appears to be instantly indexed by Google, which surprises me. I haven't had this happen before so I'm curious. I'd be interested to hear the experience of others.
Technical SEO | | waynekolenchuk0 -
Remove Bad Links Or Build New
Hello, After deeply assessing our back links we have come to the conclusion that we have too many links that have been devalued and also some spammy looking links.... Our next question is do we remove these bad links and start a fresh or do we just build new white hat links?? Thanks, Scott
Technical SEO | | ScottBaxterWW0 -
How do https pages affect indexing?
Our site involves e-commerce transactions that we want users to be able to complete via javascript popup/overlay boxes. in order to make the credit card form secure, we need the referring page to be secure, so we are considering making the entire site secure so all of our site links wiould be https. (PayPal works this way.) Do you think this will negatively impact whether Google and other search engines are able to index our pages?
Technical SEO | | seozeelot0