The Moz Q&A Forum

    • Forum
    • Questions
    • Users
    • Ask the Community

    Welcome to the Q&A Forum

    Browse the forum for helpful insights and fresh discussions about all things SEO.

    1. SEO and Digital Marketing Forum
    2. Categories
    3. Digital Marketing
    4. Web Design
    5. Reason for robots.txt file blocking products on category pages?

    Moz Q&A is closed.

    After more than 13 years, and tens of thousands of questions, Moz Q&A closed on 12th December 2024. Whilst we’re not completely removing the content - many posts will still be possible to view - we have locked both new posts and new replies. More details here.

    Reason for robots.txt file blocking products on category pages?

    Web Design
    7 3 2.5k
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes
    Reply
    • Reply as question
    Log in to reply
    This topic has been deleted. Only users with topic management privileges can see it.
    • Frankie-BTDublin
      Frankie-BTDublin last edited by

      Hi

      I have a website with thosands of products. On the category pages, all the products are linked to with the code “?cgid” in the URL. But “?cgid” is also blocked in the robots.txt file for some reason. So I'm thinking it's stopping all my products getting crawled by Google.

      Am I right here? Is there any reason why a website would want to limit so many URL's? I'm only here a week and the sites getting great traffic, so don't want to go breaking it!!!

      Thanks

      1 Reply Last reply Reply Quote 0
      • Frankie-BTDublin
        Frankie-BTDublin @AL123al last edited by

        Thanks again AL123al!

        I would be concerned about my internal linking because of this problem. I've always wanted to keep important pages within 3 clicks of the Homepage. My worry here is that while these products can get clicked by a user within 3 clicks of the Homepage, they're blocked to Googlebot.

        So the product URLS are only getting crawled in the sitemap, which would be hugely ineffcient? So I think I have to decide whether opening up these pages will improve my linking structure for Google to crawl the product pages, but is that important than increasing the amount of pages it's able to crawl and wasting crawl budget?

        1 Reply Last reply Reply Quote 0
        • AL123al
          AL123al last edited by

          Hello,

          The canonical product URLS will be getting crawled just fine as they are not blocked in the robots.txt. Without understanding your problem completely, I think the guys before you were trying to stop all the duplicate URLS with parameters being crawled and just leaving Google to crawl the canonicals - which is what you want.

          If you remove the parameter from robots.txt then Google will crawl everything including the parameter URLS. This will waste crawl budget. So better that Google is only crawling the canonicals.

          Regarding the sitemap, being present on the sitemap will help Googlebot decide what to prioritise crawling but won't stop it finding other URLS if there is good internal linking.

          Frankie-BTDublin 1 Reply Last reply Reply Quote 0
          • Frankie-BTDublin
            Frankie-BTDublin @AL123al last edited by

            Thanks AL123al! The base URL's (www.example.com/product-category/ladies-shoes) do seem to be getting crawled here & there, and some are ranking which is great. But I think the only place they can get crawled is the sitemap, which has has over 28,000 URLs on one page (another thing I need to fix)!

            So if Googlebot gets to the parameter URL through category pages (www.example.com/product-category/ladies-shoes?cgid...) and sees it's blocked, I'm guessing it can't see it's important to us (from the website hierarchy) or the canonical tag, so I'm presuming it's seriously damaging or power in getting products ranked 😕

            In Screaming Frog, 112,000 get crawled and 68% are blocked by robots. 17,000 are URL's which contain "?cgid", which I don't think is too big for Googlebot to crawl, the websites has a pretty good authority so I think we have a pretty deep crawl.

            So I suppose what really want to know is will removing "?cgid" from the robots file really damage the site? I my opinion, I think it'll really help

            1 Reply Last reply Reply Quote 0
            • AL123al
              AL123al last edited by

              This looks like the products are being appended by a parameter ?cgid - there may be other stuff attached to the end of each URL like this below:

              e.g. www.example.com/product-category/ladies-shoes?cgid-product=19&controller=product  etc

              but canonical URL is www.example.com/product-category/ladies-shoes

              These products may have had a canonical to the base URL which means that there won't be any problem with duplicates being indexed. So all well and good.

              Except.....Google has to crawl each of these parameter URLs to find the canonical. In a huge website this means that crawl budget is being consumed by unnecessary crawling of these parameterised URLs.

              You can tell Google not to crawl the parameter URLs in search console (at least in the old version you can). But you can also stop Google crawling these URLS unnecessarily by blocking them in robots txt if you are sure that the parameters are not changing how the page is looking in search.

              So long story short is that is why you may see that the URLS with parameters are being blocked in robots.txt. The canonical version URLS will be getting crawled just fine since they don't have any parameters and hence not being blocked.

              Hope that makes sense?

              Frankie-BTDublin 1 Reply Last reply Reply Quote 0
              • Frankie-BTDublin
                Frankie-BTDublin @jacobmartinnn last edited by

                Yes, it's in the robot.txt, that's the problem. Someone had to physically put it in there, but I've no idea why they would.

                1 Reply Last reply Reply Quote 0
                • jacobmartinnn
                  jacobmartinnn last edited by

                  Did you check  your robot txt file? Or check if any plugin creating this problem.

                  Frankie-BTDublin 1 Reply Last reply Reply Quote -1
                  • 1 / 1
                  • First post
                    Last post

                  Got a burning SEO question?

                  Subscribe to Moz Pro to gain full access to Q&A, answer questions, and ask your own.


                  Start my free trial


                  Explore more categories

                  • Moz Tools

                    Chat with the community about the Moz tools.

                    Getting Started
                    Moz Pro
                    Moz Local
                    Moz Bar
                    API
                    What's New

                  • SEO Tactics

                    Discuss the SEO process with fellow marketers

                    Content Development
                    Competitive Research
                    Keyword Research
                    Link Building
                    On-Page Optimization
                    Technical SEO
                    Reporting & Analytics
                    Intermediate & Advanced SEO
                    Image & Video Optimization
                    International SEO
                    Local SEO

                  • Community

                    Discuss industry events, jobs, and news!

                    Moz Blog
                    Moz News
                    Industry News
                    Jobs and Opportunities
                    SEO Learn Center
                    Whiteboard Friday

                  • Digital Marketing

                    Chat about tactics outside of SEO

                    Affiliate Marketing
                    Branding
                    Conversion Rate Optimization
                    Web Design
                    Paid Search Marketing
                    Social Media

                  • Research & Trends

                    Dive into research and trends in the search industry.

                    SERP Trends
                    Search Behavior
                    Algorithm Updates
                    White Hat / Black Hat SEO
                    Other SEO Tools

                  • Support

                    Connect on product support and feature requests.

                    Product Support
                    Feature Requests
                    Participate in User Research

                  • See all categories

                  • Body of text on category pages
                    Cantonteaco
                    Cantonteaco
                    0
                    3
                    948

                  • E-Commerce Website Architecture - Cannibalization between Product Categories and Blog Categories?
                    BeytzNet
                    BeytzNet
                    0
                    7
                    1.7k

                  Get started with Moz Pro!

                  Unlock the power of advanced SEO tools and data-driven insights.

                  Start my free trial
                  Products
                  • Moz Pro
                  • Moz Local
                  • Moz API
                  • Moz Data
                  • STAT
                  • Product Updates
                  Moz Solutions
                  • SMB Solutions
                  • Agency Solutions
                  • Enterprise Solutions
                  • Digital Marketers
                  Free SEO Tools
                  • Domain Authority Checker
                  • Link Explorer
                  • Keyword Explorer
                  • Competitive Research
                  • Brand Authority Checker
                  • Local Citation Checker
                  • MozBar Extension
                  • MozCast
                  Resources
                  • Blog
                  • SEO Learning Center
                  • Help Hub
                  • Beginner's Guide to SEO
                  • How-to Guides
                  • Moz Academy
                  • API Docs
                  About Moz
                  • About
                  • Team
                  • Careers
                  • Contact
                  Why Moz
                  • Case Studies
                  • Testimonials
                  Get Involved
                  • Become an Affiliate
                  • MozCon
                  • Webinars
                  • Practical Marketer Series
                  • MozPod
                  Connect with us

                  Contact the Help team

                  Join our newsletter
                  Moz logo
                  © 2021 - 2026 SEOMoz, Inc., a Ziff Davis company. All rights reserved. Moz is a registered trademark of SEOMoz, Inc.
                  • Accessibility
                  • Terms of Use
                  • Privacy

                  Looks like your connection to Moz was lost, please wait while we try to reconnect.