Regular Expressions for Filtering BOT Traffic?
-
I've set up a filter to remove bot traffic from Analytics. I relied on regular expressions posted in an article that eliminates what appears to be most of them.
However, there are other bots I would like to filter but I'm having a hard time determining the regular expressions for them.
How do I determine what the regular expression is for additional bots so I can apply them to the filter?
I read an Analytics "how to" but its over my head and I'm hoping for some "dumbed down" guidance.
-
No problem, feel free to reach out if you have any other RegEx related questions.
Regards,
Chris
-
I will definitely do that for Rackspace bots, Chris.
Thank you for taking the time to walk me through this and tweak my filter.
I'll give the site you posted a visit.
-
If you copy and paste my RegEx, it will filter out the rackspace bots. If you want to learn more about Regular Expressions, here is a site that explains them very well, though it may not be quite kindergarten speak.
-
Crap.
Well, I guess the vernacular is what I need to know.
Knowing what to put where is the trick isn't it? Is there a dummies guide somewhere that spells this out in kindergarten speak?
I could really see myself botching this filtering business.
-
Not unless there's a . after the word servers in the name. The . is escaping the . at the end of stumbleupon inc.
-
Does it need the . before the )
-
Ok, try this:
^(microsoft corp|inktomi corporation|yahoo! inc.|google inc.|stumbleupon inc.|rackspace cloud servers)$|gomez
Just added rackspace as another match, it should work if the name is exactly right.
Hope this helps,
Chris
-
Agreed! That's why I suggest using it in combination with the variables you mentioned above.
-
rackspace cloud servers
Maybe my problem is I'm not looking in the right place.
I'm in audience>technology>network and the column shows "service provider."
-
How is it titled in the ISP report exactly?
-
For example,
Since I implemented the filter four days ago, rackspace cloud servers have visited my site 848 times, , visited 1 page each time, spent 0 seconds on the page and bounced 100% of the time.
What is the reg expression for rackspace?
-
Time on page can be a tricky one because sometimes actual visits can record 00:00:00 due to the way it is measured. I'd recommend using other factors like the ones I mentioned above.
-
"...a combination of operating system, location, and some other factors can do the trick."
Yep, combined with those, look for "Avg. Time on Page = 00:00:00"
-
Ok, can you provide some information on the bots that are getting through this that you want to sort out? If they are able to be filtered through the ISP organization as the ones in your current RegEx, you can simply add them to the list: (microsoft corp| ... ... |stumbleupon inc.|ispnamefromyourbots|ispname2|etc.)$|gomez
Otherwise, you might need to get creative and find another way to isolate them (a combination of operating system, location, and some other factors can do the trick). When adding to the list, make sure to escape special characters like . or / by using a \ before them, or else your RegEx will fail.
-
Sure. Here's the post for filtering the bots.
Here's the reg x posted: ^(microsoft corp|inktomi corporation|yahoo! inc.|google inc.|stumbleupon inc.)$|gomez
-
If you give me an idea of how you are isolating the bots I might be able to help come up with a RegEx for you. What is the RegEx you have in place to sort out the other bots?
Regards,
Chris
Got a burning SEO question?
Subscribe to Moz Pro to gain full access to Q&A, answer questions, and ask your own.
Browse Questions
Explore more categories
-
Moz Tools
Chat with the community about the Moz tools.
-
SEO Tactics
Discuss the SEO process with fellow marketers
-
Community
Discuss industry events, jobs, and news!
-
Digital Marketing
Chat about tactics outside of SEO
-
Research & Trends
Dive into research and trends in the search industry.
-
Support
Connect on product support and feature requests.
Related Questions
-
Unexplainable drop in traffic
Hello Mozzers, I am new at Moz and certainly hope you can help me!
Intermediate & Advanced SEO | | svsanchez75
I used to consider myself very knowledgeable in SEO since I had hundreds or maybe thousands of keywords very well ranked on Google Guatemala (www.google.com.gt) for more than a decade. That was until a year ago: November 8th, 2019... On that date, all of my rankings plummeted, never to return. At first I thought it was just a hickup from the Google Algorithm, but as days turned into weeks I became very worried, so I started doing a bunch of things: Reworked the pages (except home page) to make them responsive (I used to have two versions of each page, one for desktop and one for mobile). Removed the ads from the pages to make them faster, thinking it could be due to the speed (even if my competitors sites were slower!) Fixed thousands of 404s Disavowed thousands of bad domains and spammy URLs (I never bought a single link but there were thousands of links from forums, theglobe network, etc) Removed more than 12,000 members from my own forum who had never posted anything: about 4000 of them just had created a profile to include a link to different extenral sites. Fixed a few other technical aspects... Nothing helped. In fact, my rankings have kept going down. My content is good and unique, my site has good DA but still I get outranked and buried by several sites which don't have as much information as I have. There are NO manual actions against my site according to Search Console. As I mentioned before, I have never bought any backlinks, so all of my links should be natural (although there were thousands of links to my forum from spammy sites which I disavowed). So, I am frustrated as I really don't know what the problem is. I am giving you 3 examples of Keywords and URLs of pages that were number one for years, and now are not even on first page, so that you can see them and tell me your thoughts about what may be happening: KEYWORD: VOLCANES DE GUATEMALA
URL: https://www.deguate.com/geografia/volcanes/Los-volcanes-de-Guatemala.shtml Note: was #1, now is on 5th page. Removed all the ads to see if it would help. KEYWORD: MINISTERIOS DE GUATEMALA URL: https://www.deguate.com/artman/publish/politica_ministerios/Los-ministerios-de-Guatemala.shtml Note: was #1, now is on 2nd page. Removed all the ads to see if it would help. KEYWORD: LEYENDAS DE GUATEMALA URL: https://www.deguate.com/artman/publish/misterios-leyendas/las-leyendas-mas-famosas-en-guatemala.shtml Note: Was #1, now is on 3rd page. It's stuffed with ads, as it doesn't seem to matter wheter my pages have ads or not, and since I lost my rankings on thousands of pages at least I can probably generate a little more income like that. Thank you so much for the help you can provide!0 -
Moved brand's shop to a new domain. will our organic traffic recuperate?
Hello, We are a healthcare company with a strong domain authority and several thousand pages of service related content at brand.com. We've been operating an ancillary ecommerce store that sells related 3rd party products at brand.com/shop for a little over a year. We recently invested in a platform upgrade and moved our site to a new domain, brandshop.com. We implemented page-level 301 redirects including all category pages, product detail pages, canonical and non-canonical URLs, etc.. which the understanding that there would not be any loss in page rank. What we're seeing over the last 2 months is an initial dive in organic traffic, followed by a ramp-up period of if impressions (but not position) in the following weeks, another drop and we've steady at this low for the last 2 weeks. Another area that might have hurt us, the 301 redirects were implemented correctly immediately post launch (on a wednesday), but it was discovered on the following Monday that our .htaccess file had reverted to an old version without the redirect rules. For 3-4 days, all traffic was being redirected from brand.com/shop/url to brandshop.com/badurl. Can we expect to recover our organic traffic giving the launch screw up with the .htaccess file, or is it more of an issue with us separating from the brand.com domain? Thanks,
Intermediate & Advanced SEO | | eugene_p
Eugene0 -
How Can I Rank My Website Quickly and get traffic 20k per months
Hello moz webmasters, PLZ tell me How Can I Rank My Website Quickly and get traffic 20k per months. if you have backlinks lists of edu and gov sites plz donate me. check my site https://www.steemseo.com [Link removed by a forum moderator.]
Intermediate & Advanced SEO | | tushartosi0 -
New site. How important is traffic for a new site? And what about domain age?
Hi guys. I've been building a new site because i've seen a real SEO opportunity out there. I'm a mixing professional by trade and so I wanted to take advantage of SEO to help gain more work. Here's the site: www.signalchainstudios.co.uk I'm curious about domain age. This site fairly well optimised for my keywords, and my site got pretty good content on it (i think so anyway). But it's no where to be seen on the SERP's (link at all). Is this just a domain age issue? I'd have though it might be in the top 50 because my site's services are not hard to rank for at all! Also what about traffic? Does Google want to see an 'active' site before it considers 'promoting' it up the ranks? Or are back links and good content the main factor in the equation? Thanks in advance. I love this community to bits 🙂 Isaac.
Intermediate & Advanced SEO | | isaac6631 -
Emergency duplicate of website due to DNS failure - how to minimise loss of search engine traffic?
Hi, Our client has had a disaster with their domain name registrar, where the DNS settings have been reset and it looks like the registrar won't be able to re-instate the DNS settings for four days time. This is a nightmare for lost business whilst the site and emails are offline. As a fallback, we've setup a copy of the client's website at an alternative domain name so that people can be directed there in the meantime via Facebook posts, etc. Is there anything you would recommend we do in the meantime to minimise the loss of traffic from search engines, and loss of reputation with Google? eg. using Google webmasters to tell Google about the change of address? Thank you.
Intermediate & Advanced SEO | | smaavie0 -
Traffic drop off and page isn't indexed
In the last couple weeks my impressiona and clicks have dropped off to about half what it used to be. I am wondering if Google is punishing me for something... I also added two new pages to my site in the first week of June and they still aren't indexed. In the past it seemed like new pages would be indexed in a couple days. Is there any way to tell if Google is unhappy with my site? WMT shows 3 server errors, 3 Access denied, and 122 not found errors. Could those not found pages be killing me? Thanks for any advise, Greg www.AntiqueBanknotes.com
Intermediate & Advanced SEO | | Banknotes0 -
Traffic has dropped off a cliff
I just spent thousands of $$ redesigning my site. Have a Wordpress site now with Yoast SEO. However, my organic traffic has dropped of a cliff. Can any of you help me figure out why? www.UltimateBasicTraining.com
Intermediate & Advanced SEO | | StreetwiseReports1 -
Why do i not receive google traffic?
over the 4-5 months i have published over 3000 unique articles which i have payed well over 10 000usd for, but i still only receive about 20 google visitors a day for that content. i uploaded the 3000 articles after i 301 redirected the old site to a a new domain (old site had 1000 articles, and at least 300visits from google a day), and all the old conetnt receives the traffic fine (301 redirect is working 100percent now and pr went from 0 to 3pr) articles are also good ranging from 400-800 words. 90 percent of them are indexed by google, most of them have been bookmarked to digg reddit etc website domain is over 10 years old - alltopics.com why google doesnt send me the traffic i deserve?
Intermediate & Advanced SEO | | rxesiv0