Rogerbot getting cheeky?
-
Hi SeoMoz,
From time to time my server crashes during Rogerbot's crawling escapades, even though I have a robots.txt file with a crawl-delay 10, now just increased to 20.
I looked at the Apache log and noticed Roger hitting me from from 4 different addresses 216.244.72.3, 72.11, 72.12 and 216.176.191.201, and most times whilst on each separate address, it was 10 seconds apart, ALL 4 addresses would hit 4 different pages simultaneously (example 2). At other times, it wasn't respecting robots.txt at all (see example 1 below).
I wouldn't call this situation 'respecting the crawl-delay' entry in robots.txt as other question answered here by you have stated. 4 simultaneous page requests within 1 sec from Rogerbot is not what should be happening IMHO.
example 1
216.244.72.12 - - [05/Sep/2012:15:54:27 +1000] "GET /store/product-info.php?mypage1.html" 200 77813
216.244.72.12 - - [05/Sep/2012:15:54:27 +1000] "GET /store/product-info.php?mypage2.html HTTP/1.1" 200 74058
216.244.72.12 - - [05/Sep/2012:15:54:28 +1000] "GET /store/product-info.php?mypage3.html HTTP/1.1" 200 69772
216.244.72.12 - - [05/Sep/2012:15:54:37 +1000] "GET /store/product-info.php?mypage4.html HTTP/1.1" 200 82441example 2
216.244.72.12 - - [05/Sep/2012:15:46:15 +1000] "GET /store/mypage1.html HTTP/1.1" 200 70209
216.244.72.11 - - [05/Sep/2012:15:46:15 +1000] "GET /store/mypage2.html HTTP/1.1" 200 82384
216.244.72.12 - - [05/Sep/2012:15:46:15 +1000] "GET /store/mypage3.html HTTP/1.1" 200 83683
216.244.72.3 - - [05/Sep/2012:15:46:15 +1000] "GET /store/mypage4.html HTTP/1.1" 200 82431
216.244.72.3 - - [05/Sep/2012:15:46:16 +1000] "GET /store/mypage5.html HTTP/1.1" 200 82855
216.176.191.201 - - [05/Sep/2012:15:46:26 +1000] "GET /store/mypage6.html HTTP/1.1" 200 75659Please advise.
-
Hi BM7,
I'm going to open up a ticket on this to have our engineers take a closer look at your site. Once we have an overall response, I'll post it here for other community members to view.
Cheers!
-
Thanks Megan for your reply,
Will give that a try and have blocked 2 addresses so you are reduced to 2 crawler sessions. These two measures should reduce the load considerably as long as Rogerbot respects the 7 second delay.
IMHO ignoring the Crawl-Delay set by the webmaster of the site you are crawling, which crawlers are supposed to respect, is wrong. I got a Google WMT nasty for being down 5 hours due to Rogerbot as it was the middle of the night so only got restarted in the morning.
Also, my site has around 600 discrete pages of which you crawl about 500, so even at the original 10 seconds crawl delay you could do my whole site in less than 1.5 hours, which is only required once a week. So in my mind that suggests there is no need to overrule my settings in robots.txt 'so he (Roger) can complete the crawl'.
Regards,
-
Hi there,
This is Megan from the SEOmoz Help Team. I'm so sorry Rogerbot is causing you grief! This actually might be happening because your crawl delay is too long, so rogerbot just ends up ignoring it so he can complete the crawl. If you set your crawl delay to a max of 7, then it should solve your problem. If you're still running into issues, though, please send us a message to help@seomoz.org and we'll check it out asap!
Cheers!
Got a burning SEO question?
Subscribe to Moz Pro to gain full access to Q&A, answer questions, and ask your own.
Browse Questions
Explore more categories
-
Moz Tools
Chat with the community about the Moz tools.
-
SEO Tactics
Discuss the SEO process with fellow marketers
-
Community
Discuss industry events, jobs, and news!
-
Digital Marketing
Chat about tactics outside of SEO
-
Research & Trends
Dive into research and trends in the search industry.
-
Support
Connect on product support and feature requests.
Related Questions
-
Getting spam Links pointing to our wrong url, what to do?
Hey Mozzers, Looking in my Google Search Console (Webmaster Tools), I'm getting links pointing to bogus pages on my website that result in a 404. What does one do so you can tell Google that it has been "fixed"? Do i just 301 it to another website? If I add it to my disavow list, does Google remove the error in my webmaster tools? Thank you!
Moz Pro | | Shawn1240 -
Getting my top keywords separated out in SEOmoz reports
I am using the standard functionality to produce weekly Moz reports. There does not seem to be a setting to show rankings of my most important keywords. It would be nice to have those high-level keywords on the first page. For example, I have 200 keywords in an account. I want the report to show on a page my 10 most important keywords. Is there a way to set up a Label to my keywords in order to product a report page just for those keywords?
Moz Pro | | clicktoshop0 -
AM I the only one getting misleading titles in OSE?
I am trying to locate directories in my competitor's links using OSE. Here is the workflow I am using: Filter all results to external sites only, group by site, linking to any page on the domain. Export results to csv. My competitor is in the web design industry. So I try to filter the titles of the pages linking to the competitor to look for titles containing directory. But when I click on the link for "Windsor Internet Web Design Hosting Ontario Canada Directory" I get a page with the title "Kitchen & Bathroom Showroom | London Ontario | Bathroom Vanity Showroom" Are the results really this misleading?? or am I doing something wrong here? Any insight or help would be greatly appreciated.
Moz Pro | | tdlabs0 -
Wher is the API id and key? I tried generating it several times but I get nothing...?
Wher is the API id and key? I tried generating it several times but I get nothing...?
Moz Pro | | breezego0 -
I'm trying to get 'tigi bed head' up most of all...
I'm 87th ish with this term and I don't know why?! crap result I know. With every other phrase I use 'cheap tigi bed head' 'buy tigi bed head online' etc etc, we are on the first page all day long, pls help this worthy cause? I am www.thehairroom.co.uk, free hair products for the best results. Thank You.
Moz Pro | | smoki6660 -
Why am I getting an access error when creating my first campaign?
The exact error message is: The change you wanted was rejected. Maybe you tried to change something you didn't have access to. I checked the Word Count and I am definitely <300 keywords. I have also made sure that the Branded Keywords are entered 1 per entry form. My cookies are clear and there should be no issue with my browser.
Moz Pro | | trufflelabs0 -
How are our competitors getting these inbound linking domains?
I'm currently managing SEO for my company's website, and I'm getting into link building for the first time. As part of the process, I'm using Open Site Explorer to see who's linking into our competitor sites, to get a better sense of what's available to us in our particular avenue of e-commerce. However, I'm finding that our competitors are getting inbound links from high-authority sites pretty far afield from selling jewelry - census.gov, parallels.com, warnerbros.com, and others. I try clicking through to these links, but each link starts a download of a file. I've seen .f4v, .7z, and .apk files listed as inbound links to our competitor. How is this happening? Again, I'm new to link building, so there may be a simple answer here, and if so I apologize for asking. However, this seems really strange to me, and a difficult situation to confront.
Moz Pro | | jozaksut0