Why doesn't Moz crawler follow robots.txt?
-
It is crawling the entire site, and there is stuff we do not want it to. Please advise.
-
Which I am ok with, but why am I getting duplicate content?
-
Yes, it doesn't tell them which pages not to crawl - just not to index them
-
It has been used correctly. The site is a Magento site and they have it built in. There are a lot of filters for products so it uses rel=canonical to tell Google which to index.
-
rel=canonical is not really an robots instruction file - rel=canonical is to help with duplicate copy where you have the same or similar pages and your telling search engines which pages is the preferred page.
If you don't want pages crawling you have to tell Search engines in the robots file
-
Hi There,
Rel=canonical tags tell robots, which page is actually to index out of many.
For SEOs, canonicalization refers to individual web pages that can be loaded from multiple URLs. This is a problem because when multiple pages have the same content but different URLs, links that are intended to go to the same page get split up among multiple URLs. This means that the popularity of the pages gets split up. Unfortunately for web developers, this happens far too often because the default settings for web servers create this problem.
https://moz.com/learn/seo/canonicalization
I feel you have not used it correctly, check the above article and see if it helps.
Thanks,
Vijay
-
So I made a mistake it isn't the robots.txt that is the issue. I am getting hit with a ton of duplicate content penalties so I figured that was it. The problem is that I have pages with rel=canonical tags that it is ignoring. Does Roger not read those?
-
Hi
Have to agree with the above, Rogerbot does listen to robot.txt file, unlike Bing - while they are getting better Bing ignores the robots.txt file frequently.
Ive analysed quite a few server logs over the years and Roger has always listened to the file - its usually a mistake the in the robots file.
There is an option to test your robots.txt file in GCS - while this is testing to see if Google will crawl the page - usually Roger has the same instructions as Google.
However if you are still pretty certain that Roger is ignoring robots.txt please DM your Server Logs and your website and I will take a look and analyse it for you (free of course).
Thanks
Andy
-
All major search engines, including Moz's crawler Rogerbot and Internet Archives, respect Robots.txt as a standard “robots exclusion protocol” to communicate with web crawlers and web robots.
In case you wish to exclude some specific information from all Search Engines, you can use the following sample code as reference to block specific directories.
User-agent: *
Disallow: /cgi-bin/
Disallow: /tmp/
Disallow: /junk/However, if you want to specifically block Mz's Rogerbot from crawling specific sections of your website. You may take the following reference code to block specific areas / directories in your website from rogerbot:
User-agent: Rogerbot
Disallow: /cgi-bin/
Disallow: /tmp/
Disallow: /junk/I hope this helps, If you have specific questions, please feel free to respond, I will be happy to answer them.
Regards,
Vijay
-
Hi there! Moz's crawler, rogerbot, does follow robots.txt. When he's not following robots.txt, it's usually because the robots.txt protocol is formatted improperly. Learn more about formatting your page here: https://moz.com/learn/seo/robotstxt
For more information on Roger, including how to block him, head here: https://moz.com/help/guides/moz-procedures/what-is-rogerbot
And if you want to test your formatting, try the Robots Checker here: https://support.google.com/webmasters/answer/6062598
If you're still unable to determine why rogerbot is crawling your site, feel free to write in to help@moz.com!
Got a burning SEO question?
Subscribe to Moz Pro to gain full access to Q&A, answer questions, and ask your own.
Browse Questions
Explore more categories
-
Moz Tools
Chat with the community about the Moz tools.
-
SEO Tactics
Discuss the SEO process with fellow marketers
-
Community
Discuss industry events, jobs, and news!
-
Digital Marketing
Chat about tactics outside of SEO
-
Research & Trends
Dive into research and trends in the search industry.
-
Support
Connect on product support and feature requests.
Related Questions
-
Unsolved Why Links Overview is showing '0' "Internal, followed links" when we have many?
Hi, Our site is showing '0' internal links on Links Overview when all other metrics appear to be updating correctly. Any idea why might this be happening? Or are there any issues with the MOZ getting the data from our site when it does a crawl?
Link Explorer | | Chris_Mc0 -
How is Moz DA affected by spam links? Disavow file?
So it does not appear that moz let's you upload your disavow file. So when moz calculates your DA how do spammy links factor in? After digging through our GA it appears our site was hit with the 2016 penguin update and never recovered. Our weekly visitors were 2k, then dropped to 500 and have stayed close to that level for a while. We've used the disavow tool, without success over the past 3 years. During that time we have done link out reach and built around 10 legit good quality DA links since. But we have not recovered. At this point i'm thinking I should just remove the disavow file. Moz says our spam score for our domain is 5%.
Link Explorer | | jessicapremier0 -
OSE 'request CSV' for links blank
After doing the link check on a domain, got 40 or so links. Clicked Request CSV The file that is downloaded contains no links. The column headings are there but no content Did this twice and got same blank file. OSE is showing the 40 links
Link Explorer | | cbpayne0 -
Moz Pro Tools Inbound Links
Is there a way to get a date of when a inbound link was created from a external website. And if so is there a way to add that to the Moz Pro tools reports on a csv file.
Link Explorer | | willakawillow221 -
Inbound links not found in Moz Pro(though they exist)
These sites have many inbound links but Moz Pro finds no such links. I also used SimilarWebPro Trial to confirm the existing of such links. (these sites have IP block so it requires an VPN to access with Japan IP) osakado.cc
Link Explorer | | Yuji_m
oskd.biz Could anyone figure out why Moz can't find those links to those sites? Any tricks in place to hide those links from SeoMoz? If I pay to upgrade SimilarWebPro but Id rather stay with SeoMoz Pro. thanks in advance.0 -
In "Link Opportunities" is there a way to filter by external follow links that pass equity?
In Open Site Explorer, the "Link Opportunities" feature does not seem to have a way to filter for external follow links that pass equity. This would be very useful. Is this a feature in the pipeline or is there already a way to do this?
Link Explorer | | GlennFerrell2 -
What's the Story on Mozscape Updates?
Hey gang, As you may be aware, we were considerably late with our last index release. You have my sincere apologies for that and the apologies of the entire team. In the interest of transparency, I want to try to explain what's been going on. Since stepping down as CEO, I've been asked to take on a few roles in the company. One of those is product architect (basically the product owner) of our Big Data team, who produces the Mozscape link index. For several years, that team has been almost exclusively focused on getting us closer to a near real-time indexing system that does not have scalability issues. Mozscape is currently smaller than our major competitors, and we're also often slower. Our metrics (PA, DA, MozRank, MozTrust, Spam Score, Social Data, etc) have been the unique value we provide, but it's not enough. We need to be competitive on size and freshness. Building a raw link index (without processed metrics like PA/DA et al) is hard, but it's possible. Building a link index with those metrics is really tricky, and requires computer science knowledge and skills far beyond the scope of my understanding. That's what our team's been working on, and they've made some progress, but it's been slow, hampered by unknown unknowns, and materially hurt by a lack of experienced talent we can hire to help (we've had open job posts for years now). In the meantime, our historic Mozscape index structure keeps encountering challenges - this latest round is still somewhat unexplained (we believe there's hardware issues compounded by how the system is architected to handle large domains, but there may be other issues). The team's struggled to split time between keeping the old Mozscape running and hunkering down to finish the new system. I'm trying to help them balance things as best I can, and we're going to be putting effort toward making sure we get index releases out on time. However, to do that, we'll need to scale down size, and then rebuild back up. We think we can do this while also improving the prioritization of which links we crawl (e.g. deeper on important domains that link out, less so on deep pages that don't link anywhere) so the index overall improves. However, I don't want to minimize the risks - we may have some slow updates, some smaller indices, and some less-than-ideal data in the next one or two indices while we work to remedy this issue. I HOPE we don't, and that things actually get better immediately, but we can't promise that until the work gets finished. TL;DR - Mozscape V2 is in development and will let us as big and faster as any link index. In the meantime, current Mozscape's having issues & we're making smaller indices in an attempt to diagnose and repair. As always, thanks for your understanding, continued support, and if you have any questions, feel free to leave them below. I realize that this level of service/product quality is NOT OK, and I'm doing everything in my power to fix it.
Link Explorer | | randfish8 -
How often are Moz Domain authority and Page Authority Metrics Refreshed?
How often are the MOZ metrics of DA and PA refreshed, this is because i have noticed when i conduct an Inbound link analysis of my blog, i realize that MOZ is merely detecting around 100+ only links to what is being detected on Google Webmaster Tools, any advise?
Link Explorer | | ConnectMedia0