Any tools for scraping blogroll URLs from sites?
-
This question is entirely in the whitehat realm...
Let's say you've encountered a great blog - with a strong blogroll of 40 sites.
The 40-site blogroll is interesting to you for any number of reasons, from link building targets to simply subscribing in your feedreader. Right now, it's tedious to extract the URLs from the site. There are some "save all links" tools, but they are also messy.
Are there any good tools that will
a) allow you to grab the blogroll (only) of any site into a list of URLs (yeah, ok, it might not be perfect since some sites call it "sites I like" etc.)
b) same, but export as OPML so you can subscribe.
Thanks!
Scott
-
Not at all. I guess my feeling here is that there is a sort of untapped social graph defined by blogrolls. If it were simple to harvest them upon visiting a blog (e.g. this blogger recommends...) one could do a stumble-on-steroids approach to a niche.
-
I thought you might be able to use the outbound link scraper to grab the outbound link onto the page. Pop in your URLS of the pages you want to scrape and it will spit out our a list of those domaind and urls. You can take those urls and put them into the contact finder and it will return the contact details for those sites. Combine the two spreadsheets for an epiuc list of blogs to contact for your outreach.
This is obviously for link building rather than subscribing - sorry if I have misunderstood what you were trying to do
-
Hi Keri,
That is a very cool tool, but is overkill for this. It takes far too many steps to accomplish only part of the desired goal of grabbing all blogroll URLs (within the blogroll DIV tag) and exporting the list to a valid OMPL file or URL list.
thanks!
-
nothing I saw there would do this. It looks like it could manage to list all external links, and I suppose you could manually pick the blogroll out of it.
-
Hi there,
Well, Keris response reminded me of this question and the fact that I found a tool for scraping these kind of lists:
Here it is (with some other cool tools) , have fun:
-
Hi Scott,
I'm going through older questions. Did you ever find a tool to do what you wanted to do here?
-
One thing to look at is Outwit Hub for Firefox. It might be able to help with that. It can scrape data from a page and do a lot with it. http://www.outwit.com/products/hub/. Don't know that it meets all of your needs, but I also haven't seen a response with anything better at the moment.
-
Hey Scott,
What a great question and <sigh>I don't have the answer. I am going to back to find out what people come up with here. Surely there is someone that lurks these parts that can throw something together?</sigh>
Got a burning SEO question?
Subscribe to Moz Pro to gain full access to Q&A, answer questions, and ask your own.
Browse Questions
Explore more categories
-
Moz Tools
Chat with the community about the Moz tools.
-
SEO Tactics
Discuss the SEO process with fellow marketers
-
Community
Discuss industry events, jobs, and news!
-
Digital Marketing
Chat about tactics outside of SEO
-
Research & Trends
Dive into research and trends in the search industry.
-
Support
Connect on product support and feature requests.
Related Questions
-
Is Anyone Else Having Problems With The Ranking On Pro Tools?
After checking them from the report I was emailed, some of them seem to be incorrect, or is it something my end? To be fair the majority of them are correct, I'm just querying it.
Moz Pro | | JonathanRolande0 -
Order of urls in SEOMoz crawl report
Is there any rhyme or reason to the order of urls in the SEOMoz crawl report, or are the urls just listed in random order?
Moz Pro | | LynnMarie0 -
How can a site have a backlink from Barclays website?
Hi, I have entered a competitiors website www.my-wardrobe.com into Open Site to see who they get links from and to my surprise they have a load from Barclays Business Banking. When I visit the page I can not see the links. But if I search the pages source code for my-wardrobe, there I have it, a link to my-wardrobe.com. How have they done this? Surely Barclays haven't sold them it? And more so, why are they receiving link juice when you cant even see the link on the Barclays page in question - http://www.barclays.co.uk/BusinessBanking/P1242557952664 Thanks | |
Moz Pro | | YNWA
| | <a <span="">href</a><a <span="">="</a>http://www.my-wardrobe.com" class="popup" title="Link opens in a new window" rel='' onmousedown="dcsMultiTrack('DCS.dcsuri','BusinessBankingfromBarclays/Footer/wwwmywardrobecom', 'WT.ti', '','WT.dl','1');"> |
| | www.my-wardrobe.com |
| |
|
| | |0 -
Which of these is the best guest blogging site
Which of these is best: http://www.guestblogit.com/http://www.bloggerlinkup.com/ http://www.Myblogguest.comhttps://www.helpareporter.com/ I like the look of MyBlogguest.com so far. We are wanting to guest write articles to be published on quality sites in exchange for one or two links back (link building) Also, how do you choose what article topics to write. So far my strategy is to look at the industry's biggest sites and to use the "Top Pages" tab in OSE to look for hot topics. Thanks!
Moz Pro | | BobGW0 -
404 errors in SEOMoz crawl tool
I currently have several 404 errors in the latest crawls from SEOMoz. Here is an example of the error. http://dealerplatform.com/blog/2011/10/23/videos-for-auto-dealers/www.dealerplatform.com/ In all cases the error is a result of www.dealerplatform.com being added to the real url. Anyone seen this before? The site is a wordpress mutlisite. I don't see where this incorrect link is showing up anywhere on the website. Any advice would be helpful. Thanks
Moz Pro | | Chris_Gregory0 -
Keyword Difficulty Tool - How to use it the best way?
Hi, I am freshman, both, here on SEOmoz and in SEO generally and have a question concerning the assessment of KW difficulty. I did browse through the Q&A-Section (great content!!) but could not find a relevant answer to my problem. So I am currently building my initial keyword list, our website is only 2 months old, so we are still in the very early stage. Fashion is a very competitive area, so the KW Diff. Tool indicates high difficulty for a lot of words and phrases. however, I identified some with percentages <50 or even lower than that. Then I compared the results of the SEOmoz tool to Google Adwards and the Google results for competitiveness differed significantly. For example: for the KW Personal Shopping I got KWD from Seomoz 33% (in Germany) and from Google 0,5 for broad and 0,64 for exact search. I am quite confused how to make the right choices for my KWs now. Which metrics should I consider most? What else apart from the competition factor is behind the metric KW diff.? Does it matter in any way that I search from Germany for Germany in German? Do you have any further recommendations for the process of identifying the best Kws? Thanks a lot in advance, best from Berlin Tani
Moz Pro | | TaniBogi0 -
Internal links not showing in Open Site Explorer
So I'm working on a law firm site and looking at the links for pages in OSE. For practice areas, the links to each practice area are in the left hand menu on every page of the site. Can anyone help me with this question: Example: http://www.comitzlaw.com/personal-injury/car-accidents.html When I plug this URL into OSE, it only shows one linking page, www.comitzlaw.com/practice-areas.html, yet there is a link to this on every other page in the site. When I plug in a random competitors page, www.lesagelblaw.com/Personal-Injury-Overview/Car-Accidents.shtml, it does show all the internal pages linking to it. Since I'm not using a flash menu or javascript, any ideas as to why no internal links are showing up in OSE? Even when I plug in the main URL for the home page, it only shows 4 other internal pages linking to it, yet there is a link on every page. What am I doing wrong?
Moz Pro | | c2g0 -
Recent backlinks in Open Site Explorer as not showing
I saw the note today that the link index did it's monthly update, yet our site www.oznappies.com still only shows 1 linking root domain and I know there are many more links now. What do I need to do to get open Site Explorer to use the latest data? I enter our site, create the report and only see old information from 6 may 2011's link index. I have the same issue with competive link finder, links I know we have on the sites listed for our competitors are not showing for our site.
Moz Pro | | oznappies0