What Sources to use to compile an as comprehensive list of pages indexed in Google?

sp80

As part of a Panda recovery initiative we are trying to get an as comprehensive list of currently URLs indexed by Google as possible.

Using the site:domain.com operator Google displays that approximately 21k pages are indexed. Scraping the results however ends after the listing of 240 links.

Are there any other sources we could be using to make the list more comprehensive? To be clear, we are not looking for external crawlers like the SEOmoz crawl tool but sources that would be confidently allow us to determine a list of URLs currently hold in the Google index.

Thank you /Thomas

Dr-Pete

We don't usually take private info in public questions, but if you want to, Private Message me the domain (via my profile). I'm really curious about (1) and I'd love to take a peek.

sp80

Thanks Pete,

As always very much appreciate your input.

1/ We aren't using any parameters and when using the filter=0 we are getting the same results. For my just done test I was only able to pull 350 pages out of 18.5k pages using the web interface. If anyone has any other thoughts on this please let me now.

2/ That is a great idea. Most of our pages live in the root directory to keep the URL slugs short so unfortunately this one will not help us.

3/ Another good idea. I understand this approach is helpful to see your coverage of wanted pages in the Google index but won't be able to help you determine superfluous pages currently in the Google index unless I misunderstood you?

4/ We are using ScreamingFrog and I agree its a fantastic tool. The index size with ScreamingFrog is showing not more than 300 pages which is our final goal.

Overall we are seeing continuous yet small drops to the index size using our approach of returning 410 response codes for unwanted pages and dedicated sitemaps to speed up delisting. See http://www.seomoz.org/q/panda-recovery-what-is-the-best-way-to-shrink-your-index-and-make-google-aware

We are just trying to get a more complete list of whats currently in the index to speed up delisting.

Thank you for your reference to the Panda post I remember reading it before and will give it another go right now.

One final question, in your experience dealing with Panda penalties, have you seen scenarios where it seems the delisting/penalizing of a site has only happened for a particular CCTLD of google or just the homepage? See http://www.seomoz.org/q/panda-penguin-penalty-not-global-but-only-firea-for-specific-google-cctlds It is what we are currently experiencing and trying to see if other people have observed something similar.

Best /Thomas

Dr-Pete

If you're willing to piece together multiple sources, I can definitely give you some starting points:

(1) First, dropping from 21K pages indexed in Google to 240 definitely seems odd. Are you hitting omitted results? You may have to shut off filtering in the URL (&filter=0).

(2) You can also divide the site up logically and run "site:" on sub-folders, parameters, etc. Say, for example:

site:example.com/blog

site:example.com/shop

site:example.com/uk

As long as there's some logical structure, you can use it to break the index request down into smaller chunks. Don't forget to use inurl: for URL parameters (filters, pagination, etc.).

(3) This takes a while, but split up your XML sitemaps into logical clusters - say, one for major pages, one for top-level topics/categories, one for sub-categories, one for products. That way, you'll get a cleaner could of what kind of pages are indexed, and you'll know where your gaps are.

(4) Run a desktop crawler on the site, like Xenu or Screaming Frog (Xenu is free, but PC only and harder to use. Screaming Frog has a yearly fee, but it's an excellent tool). This won't necessarily tell you what Google has indexed, but it will help you see how your site is being crawled and where problems are occurring.

I wrote a mega-post a while back on all the different kinds of duplicate content. Sometimes, just seeing examples can help you catch a problem you might be having. It's at:

http://www.seomoz.org/blog/duplicate-content-in-a-post-panda-world

sp80

Does anyone have any insight on this? If the answer is simply there is no better approach than look at the limited data available through the Google UI this would be helpful as well.

Welcome to the Q&A Forum

Browse the forum for helpful insights and fresh discussions about all things SEO.

What Sources to use to compile an as comprehensive list of pages indexed in Google?

Got a burning SEO question?

Browse Questions

Explore more categories

Related Questions

What to do if lots of backend pages have been indexed by Google erroneously?

Google Not Indexing App Content

My blog is indexing only the archive and category pages

Can Google index PDFs with flash?

Should I use rel=canonical on similar product pages.

WordPress redesign: using posts as pages?

How long does google take to show the results in SERP once the pages are indexed ?

How do I index these parameter generated pages?