Abnormal crawl issues appearing in my Moz results
-
I have been asked to look at a site for a friend and was more than surprised to see 16,9k crawl issues appear in the dashboard... of this 6,238 are duplicate page content and 5878 are duplicated page titles.
What on earth is going on? I have spoken to the web developer as it appears there is a dev site somewhere and this is his response
[Can I stress that Google determines which site was in the index first and then removes other sites it sees as having duplicate content. Our dev sites appearing in the search index would not affect your ranking due to duplicate content as Google would see your site as the first site with the content]
As I cannot make contact with him, I am scratching my head, surely a dev site should be no-indexed, it sounds as though he is saying that its ok because Google will take the main site as the first site with the content...
Very confused! Help need MOZ community.
Manythanks,
Sarah
-
Thanks again Dirk. I like your direct and knowledgeable responses. I have sent a Linkedin connection!!
Many thanks,
Sarah
-
Hi Sarah,
Googlebot will follow these links as well and discover these "useless" pages (the are off course not useless from human perspective but they don't add value for bots - and they will be considered as duplicates). Duplicates are no reason for "punishment" - so you could just let them be. Personally I would put a nofollow on these links or add a "noindex" tag to the login page. Normally you shouldn't use nofollow on internal links - but login pages are an exemption on this (check also https://searchenginewatch.com/sew/news/2298312/matt-cutts-you-dont-have-to-nofollow-internal-links : "Of course, there are always exceptions to the rule, and things like login pages can be the exception. He said it doesn’t hurt to put the nofollow link for a link pointing to a login page, or things like terms and conditions or other “useless” pages. However, it doesn’t hurt at all for those pages to be crawled by Google."
For the practical part - if you add an additional question to a question which has been marked as answered - only the ones who have already answered will see the additional question. To be on the safe side - it's better open a new question if you want other people to have a look at it.
Hope this helps,
Dirk
-
hello Dirk, thank you for that great answer, we have since been doing a bit more digging of our own and before we go back to the web developer we want to check what should be happening with the links the we are finding duplicated as we are seeing that the issues relating to Duplicate Pages are coming from links from the login page which shows information about where the user was redirected from.
For example, if the visitor is not logged on and wishes to wish-list an item, they will be redirected to the login page, with the item code and intended action in the url; which can then continue on to the desired page once logged on.
The MOZ crawler is seeing these pages as having Duplicated Content whilst they are all the same apart from a piece of information in the URL. Should we be blocking these duplications? Are they a risk to us? What should we be doing?
I have also added this as a new question - I am quite new to this community thing so wasn't sure which was the best way to ask the question.
Many thanks again,
Sarah
-
Moz is only indexing pages it's crawler is able to find. This implies that on your production site you have links to your development site.
Don't really agree with what your dev is saying - he should correct these links first; put a noindex on these pages. Alternative - put a password on the dev site so it's only accessible with a password. If a lot of users are putting links to your dev site it could become more important than your main site. Google will try to choose the most appropriate site - but you have no guarantee that it will choose the right version. In any case - that's not the type of risk you should be willing to take.
Once this is done - you can request a removal of these pages via the search console.
If all pages are removed from the index you can adapt the robots.txt to prevent access to the Google & other bots. Do this only after all pages are removed - if not Google will never find the noindex directive.
Dirk
Got a burning SEO question?
Subscribe to Moz Pro to gain full access to Q&A, answer questions, and ask your own.
Browse Questions
Explore more categories
-
Moz Tools
Chat with the community about the Moz tools.
-
SEO Tactics
Discuss the SEO process with fellow marketers
-
Community
Discuss industry events, jobs, and news!
-
Digital Marketing
Chat about tactics outside of SEO
-
Research & Trends
Dive into research and trends in the search industry.
-
Support
Connect on product support and feature requests.
Related Questions
-
Unsolved How to concat moz.com
Hello,
Moz Pro | | GrzegorzZ
I wanted to contact moz.com. I started my 30 day trial to test service.
After 2-3 days totally forgot about it. I remembered about moz.com when i received invoice saying they're gonna charge me.
I immediately wrote email to them that I do not want this service. I forgot and only put credit card data because i was required to. And would like a refund since it was 2-3 hours after bill. Unfortunately sending messages via their contact form is not an option. There is no confirmation on my email they received message nor return message from them.0 -
Setting up MOZ to run on Staging
Hi Moz, We would like to setup Moz to run on our Staging Server. This would be extremely valuable as it alerts us to new SEO issues/risks in a controlled and secure environment that is not exposed to production. Our internal team has recommended potentially setting up a reverse proxy server that will validate either via Moz's/Rogerbot's http header or IP and allow Moz access to our Staging environment to crawl. Is this something that we can setup with Moz? Are there other ideas to enable Moz to crawl our Staging server?
Moz Pro | | kriskunisch0 -
Canonicalisation/ Moz/ WP.
Hi, I've recently put a new site live using Wordpress, and have hooked it up to Moz Analytics. Moz is reporting >1000 pages with 'missing meta description page', and >250 with 'duplicate page content'. The issues seem to be being flagged by pages with '?sort=asc&view=main', ?show=paged&view=main, '?sort=ratedup&show=all&view=list', etc. I have a number of category pages, which users can filter by etc to give them the information in the way they want to see it. I have this ' I haven't got many pages indexed by Google yet, so can't see if they're hitting the same issues as Moz. Any help/ advice appreciated.
Moz Pro | | newstd1000 -
On page analysis showing old results
Hi, My on page crawl analysis is showing that I have 38 duplicate content issues. These issues are because we had used tags in our blog. We remove all the tags 10+ day ago but the crawl that was done 5 days ago is still showing the tags as causing a duplicate content issue. Why would this be? We cannot correct is as it is already corrected. How do we get moz to crawl the correctly? Thanks Andrew
Moz Pro | | Studio330 -
1 page crawled - again
Just had to let you know that it happend again. So right now we are at 2 out of the last 4 crawls. Uptime here is 99,8% for the last 30 days, with a small downtime due to an update process at the 18/5 from around 2:30 to 4:30 GMT In relation to: http://moz.com/community/q/1-page-crawled-and-other-errors
Moz Pro | | alsvik0 -
Only One page crawled..Need help
I have run a website in Seomoz which have many URLs with it. But when I saw the seomoz report that showing Pages Crawled: 1. Why this is happen my campaign limit is OK Tell me what to do for all page crawling in seomoz report. wV6fMWx
Moz Pro | | lucidsoftech0 -
How do I force a crawl?
In the campaign overview it reads that 0 pages were crawled. Also got an email saying that a comprehensive audit will be done in 7 days. But the 'crawl in progress' wheel disappeared. I think it stopped, and I need to submit that report to substantiate buying the tool! How do I force a crawl?
Moz Pro | | ilhaam0 -
SEOMoz Q&A having some issues?
Apologies if this has been asked. I was browsing Q&A and got a Ruby error page, reloaded and everything was fine. However a bunch of the topics are kind of messed up, and it looks like a chunk of them are missing now. For example my 'questions I've answered' section of My Q&A only has the newest question and the oldest question, and when I sort the SEOMoz Tools category by newest I get one recent post and then the next one is from November of last year. You guys having database problems?
Moz Pro | | icecarats0