Google has deindexed 40% of my site because it's having problems crawling it

Bajram.Kurtishaj

Hi

Last week i got my fifth email saying 'Google can't access your site'. The first one i got in early November. Since then my site has gone from almost 80k pages indexed to less than 45k pages and the number is lowering even though we post daily about 100 new articles (it's a online newspaper).

The site i'm talking about is http://www.gazetaexpress.com/

We have to deal with DDoS attacks most of the time, so our server guy has implemented a firewall to protect the site from these attacks. We suspect that it's the firewall that is blocking google bots to crawl and index our site. But then things get more interesting, some parts of the site are being crawled regularly and some others not at all. If the firewall was to stop google bots from crawling the site, why some parts of the site are being crawled with no problems and others aren't?

In the screenshot attached to this post you will see how Google Webmasters is reporting these errors.

In this link, it says that if 'Error' status happens again you should contact Google Webmaster support because something is preventing Google to fetch the site. I used the Feedback form in Google Webmasters to report this error about two months ago but haven't heard from them. Did i use the wrong form to contact them, if yes how can i reach them and tell about my problem?

If you need more details feel free to ask. I will appreciate any help.

Thank you in advance

C43svbv.png?1

DirkC

Great news - strange that these 608 errors didn't appear while crawling the site with Screaming Frog.

Bajram.Kurtishaj

We found the problem. It was about website compression (GZIP). I found this after crawling my site with Moz, and saw lot's of pages with 608 Error code. Then i searched in Google and saw a response by Dr. Pete in another question here in Moz Q/A (http://moz.com/community/q/how-do-i-fix-608-s-please)

After we removed the GZIP, Google could crawl the site with no problems.

Bajram.Kurtishaj

Dirk

Thanks a lot for your help. Unfortunately the problem remains the same. More than 65% of site has been de-indexed and it's making our work very difficult.

I'm hoping that somebody here might have any idea of what is causing this so we can find a solution to fix it.

Thank you all for your time.

DirkC

Hi

Not sure if the indexing problem is solved now, but I did a few other checks. Most of the tools I used where able to capture the problem url without much issues even from California ip's & simulating Google bot.

I noticed that some of the pages (example http://www.gazetaexpress.com/fun/) are quite empty if you browse them without Javascript active. Navigating through the site with Javascript is extremely slow, and a lot of links don't seem to respond. When trying to go from /fun/ to /sport/ without Javascript - I got a 504 Gateway Time-out

Normally Google is now capable of indexing content by executing the javascript, but it's always better to have a non-javascript fallback that can always be indexed (http://googlewebmastercentral.blogspot.be/2014/05/understanding-web-pages-better.html) - the article states explicitly

If your web server is unable to handle the volume of crawl requests for resources, it may have a negative impact on our capability to render your pages. If you’d like to ensure that your pages can be rendered by Google, make sure your servers are able to handle crawl requests for resources.

This could be the reason for the strange errors when trying to fetch like Google.

Hope this helps,

Dirk

Bajram.Kurtishaj

Hi Dirk

Thanks a lot for your reply.

Today we turned off the firewall for a couple hours and tried to fetch the site as Google. It didn't work. The results we're the same as before.

This problem is starting to be pretty ugly since Google has started now not showing our mobile results as 'mobile-friendly' even though we have a mobile version of site, we are using rel=canonical and rel=alternate and 302 redirects for mobile users from desktop pages to mobile ones when they are browsing via smartphone.

Any other idea what might be causing this?

Thanks in advance

DirkC

Hi,

It seems that you're pages are extremely heavy to load - I did 2 tests - on your homepage & on the /moti-sot page

Your homepage needed a whopping 73sec to load (http://www.webpagetest.org/result/150312_YV_H5K/1/details/) - the moti-sot page is quicker - but 8sec is still rather high (http://www.webpagetest.org/result/150312_SK_H9M/)
I sometimes noticed a crash of the Shockwave flash plugin, but not sure if this is related to your problem;

I crawled your site with Screaming Frog, but it didn't really find any indexing problems - while you have a lot of pages very deep in your sitestructure, the bot didn't seem to have any specific troubles to access your page. Websniffer returns a normal 200 code when checking your sites - even with useragent "Google"

So I guess you're right about the firewall - may be it's blocking the ip addresses used by Google bot - do you have reporting from the firewall which traffic is blocked? Try to search for the useragent Googlebot in your logfiles and see if this traffic is rejected. The fact that some sections are indexed and others not could be related to the configuration of the firewall, and/or the ip addresses used by Google bot to check your site (the bot is not always using the same ip address)

Hope this helps,

Dirk

Welcome to the Q&A Forum

Browse the forum for helpful insights and fresh discussions about all things SEO.

Google has deindexed 40% of my site because it's having problems crawling it

Got a burning SEO question?

Browse Questions

Explore more categories

Related Questions

How to fix Google index after fixing site infected with malware.

New Website, New URL, New Content - What do we do with the old site? Are 301's the only option?

Javascript to manipulate Google's bounce rate and time on site?

Intuit's Homestead web developer

Https-pages still in the SERP's

About Bot's IP

What's the SEO impact of url suffixes?

Switching Site to a Domain Name that's in Use