Does google scrape links from PDF files? do these links pass link juice?
-
Title is pretty much the whole question.
-
I made a test and it seems that yes, the links from pdf count for ranking.
The test is on my Romanian blog http://seogan.ro/link-building-pdf-urile-o-sursa-de-linkuri-test
You can find an English translation here: http://www.seogan.com/pdf-link-building
Hope it helps.
-
Yes it does according to Google tech spec http://code.google.com/apis/searchappliance/documentation/50/admin_crawl/Introduction.html
which specifically states if follows html links in pdf 'It follows HTML links in PDF files, Word documents, and Shockwave documents'. Google's own api docs carry more weight than a comment in a forum_._ If they are licencing this out as an application it would suggest that the same technology is available in the main engine as does Dunamis's comment about a listing in a pdf document being found in search results.
You can test for youself by publishing a pdf with a link to a info page that does not show up in any other links. Include the pdf in your sitemap but not the test page and check if it shows in googles index site:yoursite.com the next time it crawls.
This also gives some insight in an interview with Matt Cutts - http://www.stonetemple.com/articles/interview-matt-cutts-012510.shtml
Eric Enge: What about PDF files?
Matt Cutts: We absolutely do process PDF files. I am not going to talk about whether links in PDF files pass PageRank. But, a good way to think about PDFs is that they are kind of like Flash in that they aren't a file format that's inherent and native to the web, but they can be very useful. In the same way that we try to find useful content within a Flash file, we try to find the useful content within a PDF file. At the same time, users don't always like being sent to a PDF. If you can make your content in a Web-Native format, such as pure HTML, that's often a little more useful to users than just a pure PDF file.
-
This person seems to think no: http://www.google.fr/support/forum/p/Webmasters/thread?tid=14c5fe970fe84361&hl=en
but i'm not sure how much i can trust a random comment from a random source. any evidence for either argument?
EDIT: And this person seems to think they do pass link juice: http://www.whydowork.com/blog/link-building/274/
Could a mod remove the marked as answered? i don't think i am able to remove it, and the question isn't really answered.
-
yes, but do they crawl the links they find in these documents, or do they just index their contents.
-
Hmmm although i thought you had answered my question, i actually feel that you have not... Yes the links you provided state that google scrapes pdfs and even OCRs pdfs to get a better idea what is in them, but i don't see anywhere that they mention crawling the urls they find in these pdf documents.
-
Google definitely does index the contents of pdf files. I found this out the hard way as I had a real estate pdf on my site that I wanted to have listed in the index, but I didn't know that the contents would be crawled. The pdf contained some listings that I was not legally allowed to advertise on my site. (It was legal for me to give someone a report with the listings in it though).
When another realtor was searching for their own listing, my pdf came up. I got in trouble. I'm ok now though.
-
Have a look at this article http://searchenginewatch.com/article/2067225/Google-Does-PDF-Other-Changes it explains some of the doc library search for pdf files and Google's statement here http://googleblog.blogspot.com/2008/10/picture-of-thousand-words.html.
Got a burning SEO question?
Subscribe to Moz Pro to gain full access to Q&A, answer questions, and ask your own.
Browse Questions
Explore more categories
-
Moz Tools
Chat with the community about the Moz tools.
-
SEO Tactics
Discuss the SEO process with fellow marketers
-
Community
Discuss industry events, jobs, and news!
-
Digital Marketing
Chat about tactics outside of SEO
-
Research & Trends
Dive into research and trends in the search industry.
-
Support
Connect on product support and feature requests.
Related Questions
-
Links not showing
Hi, we have been putting links to our website www.caffeinemarketing.co.uk from other sites we are connected to. Although these sites are up and running and have been crawled, the links are not showing up on mot. What else do I need to do? Is it anything to do with whether I put http://www. or just www. Please can someone advise.
Link Building | | Caffeine_Marketing
Thanks0 -
Does a Soft 404 pass on link juice?
Hey guys, Just wondering if a Soft 404 passes on link juice? I know they technically show up as a broken backlink but since it takes them to the homepage, does the link juice get passed on or is it still considered dead and the link juice is lost?
Link Building | | bks_seo0 -
If you discount a subscription when the client links to you, does that classify as a paid link?
For example, we all subscribe to MOZ, but if we received a 25% discount on our subscription if we linked back, would that be considered by Google as a paid link, or would it be seen as a sales tool / promotion?
Link Building | | JonathanSmith0 -
Link Building - Affiliate Link Programs
I'm getting some conflicting information when researching this... I want to create the most organic link building strategy that not only gets high authority sites linking to a website, but relevant websites. I'm creating an affiliate link program where I link to partner and affiliate companies, who in turn link back to me. The question is where to put these links. I was initially thinking of putting them in the footer so they add credibility to every page, and in turn ask my affiliates to do the same thing. However; some sources say that Google no longer looks at links in the footer section. Is it true that Google doesn't count links in footer sections? And do you have any tips on creating this Affiliate Link Program?
Link Building | | reidsteven750 -
Pinterest and Link Juice
I can't understand why our Pinterest boards don't show up as external links (nor do our YouTube videos). However, one of our competitor's Pinterest boards is in the top ten as far as their external links go. Clues?
Link Building | | Greatmats0 -
Links Removed from Site but not Google Webmaster Tools
I have a new client who has been fighting to get past an unnatural links penalty. One of the biggest reasons behind it was that they paid to get a followed link/ad on a blog network, well this ad ended up being a site wide link which ended up giving them thousands of back links within a few days. After getting the owner of the blog network to remove the link i manually spot checked a handful of pages and used screaming frog to crawl looking for a link referencing their site. From what i can tell they are all gone. However they are still in Google Webmaster Tools as well as Google's index and it has been a couple months since they have been removed. Does anyone have any advice on getting them removed from Google's entities even though they are already off of the website? Thank you in advance,
Link Building | | kchandler
Kyle0 -
Root Domain Link for Affiliate's Link
It seems my affiliate link: http://www.hrmsplugins.com?partners=21 is not being considered as a "root domain" backlink when this link is used on their website. Is there a reason for this?
Link Building | | delphia0 -
External link
Hi guys! You must get many questions about external links... This is mine. If I get an external link with good authority ranking in google, but it is not relevant to the site (for example, my site is about marquees, but we could get links from schools, which are our main clients, or any type of client) Does that affect negatively somehow to the rankings? or there is good value on it even is not related? Thanks! Adriana
Link Building | | extrememarquees0