How to download an entire Website (HTML only), ready to rehost
-
Hi all,
I work for a large retail brand and we have lots of counterfeit sites ranking for our products. Our legal team seizes the websites from the owners who then setup more counterfeit sites and so forth.
As soon as we seize control of a website, the site content is deleted and subsequently it falls out of the SERPs to be immediately replaced by the next lot of counterfeit sites.
I need to be able to download a copy of the site before it is seized, so that once I have control of it I can put the content back and hopefully quickly regain the SERPs (with an additional 'counterfeit site' notice superimposed on that page in JS).
Does anyone know or can recommend good software to be able to download an entire website, so that it can be easily rehosted?
Thanks
FashionLux
(Edited title to reflect only wanting to download html, CSS and images of site. I don't want the sites to actually be functional - only appear the same to Google)
-
Thanks for the detailed explanation.
If you know of any software or techniques to crawl and download multiple (html) pages and images of a site then please let me know.
There are many programs designed to crawl websites and grab the html code. Legitimate sites are often duplicated in this manner. You can try searching a couple relevant terms or searching black hat seo sites.
-
"If it is a very basic pure html/css site, you can pretty much achieve your goal." - Yes this is exactly what I need, I don't want the site to be functional and allow users to place orders (which could happen for non-JS users who don't see the notice that fills the entire screen). I don't want to do anything apart from rehost the site and put a big message up that says "THIS SITE WAS A SCAM - BEWARE OF OTHER SCAM SITES" and cannot be closed down.
"Do you obtain control over just the domain?" - Yes only the domain, not the hosting. We go through legal proceeding to prove the site is illegally selling counterfeit goods and obtain the blank domain.
"I understand your intentions are good, but the method is not complaint with Google's Guidelines." Fair point, but Google shouldn't rank these sites in the first place - they have no genuine links and should be banned already. Google aren't spotting this, so I have to fix Google's **** up. If the site gets banned I couldn't care less. Whilst they rank they serve a genuine purpose of (a) showing users that there are counterfeit sites and they need to be wary and (b) new sites have to better the SEO ability of the old ones in order to rank on page 1.
"Your goal is purely to manipulate search engine results which makes these activities black hat and subject to penalty." Yes but it doesn't matter if the domain is banned, it's not my genuine website and has no links back to my genuine site. I'm not going to host the sites on the same server as our genuine site so no risk to the company. Really I couldn't care less if it gets banned - the counterfeit sites are ranking due to black hat techniques - its in my interest for Google to eventually work it out and fix their algo as it will stop the hundreds of other counterfeit sites from ranking too.
"You can use every social media page, etc. If you put in the time and effort, these pages will rank very well in SERPs."
Yes I could, by building links to social media pages for the hundreds of search terms currently dominated by counterfeiters but this is not a good idea for two reasons:
1. Trying to rank social media sites for irrelevant terms isn't a good thing - you wouldn't do it for users if this situation wasn't happening. As you said already, this is a form of trying to manipulate SERPs and I wouldn't want to risk these genuine SocMed pages getting banned because of this.
2. There are hundreds of search terms to optimise for, and 8 remaining slots on Google to fill for many of these. These sites are also powerful in their SEO strength - 17 counterfeit sites made it into Majestic's top 1million sites by links - these sites have literally tons of scummy, comment box spammed links pointing at them and they are ranking (shame on Google). Competing against these isn't possible via white hat methods and I'm not a black hat kind of guy.
My thought process is - Why try and compete against these sites (and waste A LOT of time and effort) trying to bump them down the rankings when they've already done the hard work of optimisation and link building for these terms? I could simply 're-use' them for a genuine purpose (making our customers beware of ordering from unofficial websites).
The previous owner won't sue us for re-using their content - that involves making themselves known to authorities and they'd get arrested in turn for their illegal activities.
I'm happy to debate it more as its an interesting subject and I don't want to waste time going down the wrong route, but I think re-using the sites is the best option - I just need to get copies of them so they LOOK the same to Google and hopefully keep their SERPs.
If you know of any software or techniques to crawl and download multiple (html) pages and images of a site then please let me know.
Thanks for all of the responses
FashionLux
-
Thanks for the response.
"you can download the the html but not the files themselves" - the html is all I need. I don't want the site to actually work so having only the html files is perfect.
I can go to the homepage and manually save it, and go through 100+ pages and manually download them - I just wanted to ask if there was any software that would do this for me and save some leg work.
Thanks again
-
Most sites are database driven. The public does not have direct access to the database. Accordingly you cannot download the full functioning website in the manner you desire.
If it is a very basic pure html/css site, you can pretty much achieve your goal.
Do you obtain control over just the domain? Or do you have access to their hosting account? If you gain access to the hosting account, you can request the host restore the site from a backup.
Even if you gain access to the full site, you really need to be careful. Your goal is purely to manipulate search engine results which makes these activities black hat and subject to penalty. I understand your intentions are good, but the method is not complaint with Google's Guidelines.
If you own the brand, and you have a trademark, you can build quality sites promoting the brand. You can use every social media page, etc. If you put in the time and effort, these pages will rank very well in SERPs.
Some great legal victories are being won in the US to help with these types of issues. Coach recently won a similar case. It's great to hear the good guys are gaining some ground.
-
Dude, you wont be able to do that, the files are stored on the server behind a password locked folder.
Like Ryan said you can download the the html but not the files themselves.
As long as you get the content that should be enough, put it into a word doc and paste it back up once you have the domain, doesn't even need a template.
You need to stop them from re-using the content on another site.
-
Hi Ryan,
Thanks for the reply. To clarify, the site is deleted prior to me gaining control of it - by the time it comes into my hands it's completely blank, so FTP'ing isn't an option.
The site owners are essentially scamming members of the public by charging hundreds of dollars for goods that are never delivered. We've seized hundreds of sites through legal proceedings, but more keep popping up the moment we get hold of them.
These sites rank for hundreds of popular search terms (some have hundreds/thousands of spammy inbound links), so bumping them off page 1 for all SERPs isn't achievable.
By seizing the sites, keeping the content, but making the site non-functioning (imagine a popup image that fills the screen and can't be escaped) it will hopefully mean we own these SERPs and new counterfeit sites have to try and outrank them.
In turn we'll seize those sites, so the next wave of counterfeit sites have to do even more link building - eventually (maybe years) they'll realise its not worth it and give up.
Manually downloading individual webpages isn't an option, so I'm wondering if theres any programmes that can download all html files for a website so I can then just upload them via ftp once the site has been seized and add my javascript image
Thanks for all of the responses
FashionLux
-
Based on your question I am not clear if the site is deleted prior to your gaining control over the site.
If you are trying to copy a site before you have control over it, all you can do is download the HTML of the various web pages. If you spend a bit more time, you may be able to figure out file names on the server and download them, but that is moving down a path of internet security and hacking.
If you are trying to copy a site after you have control over it, the easiest method to capture everything would be a cPanel backup. cPanel is the most popular software used to administrate Apache web servers. That is the most likely hosting environment for counterfeit sites. A single cPanel backup will capture everything.
Otherwise you can go through and copy the public_html folder (or whatever the main folder is called, it will vary based on server setup) along with the database and other settings you wish to retain such as e-mail.
Understand the old site owner will still have all the passwords and an understanding of the code. While it is unlikely, they could leave themselves backdoors into the site as well. This is one reason why maintaining their site is not likely to be a good idea.
Once you began running these sites from your server, what is the plan? You would place a "counterfeit" notice and then ??? that's it? Or would you redirect them to your site? If you redirect them to your site and maintain these sites up on an ongoing basis, it can be seen as a network of doorway sites.
I understand what you are doing and why. The issue is you are taking actions purely based on search engine rankings. To do such for a short period such as 30-60 days is likely fine. To do it on a more permanent basis will likely lead you to a penalty.
-
Hi Dean,
Heather is right! you should access the websites through FTP. Also if there are databases then you should be able to export the data from the software that is managing it.
Istvan
-
Hi Dean
Could you not just use your FTP client (like Filezilla or Dreamweaver) to pull the entire site content down, save it locally, ready to upload later? Or do you not have FTP details of the sites you're taking over?
Sorry if I've miss understood the question
Heather
Got a burning SEO question?
Subscribe to Moz Pro to gain full access to Q&A, answer questions, and ask your own.
Browse Questions
Explore more categories
-
Moz Tools
Chat with the community about the Moz tools.
-
SEO Tactics
Discuss the SEO process with fellow marketers
-
Community
Discuss industry events, jobs, and news!
-
Digital Marketing
Chat about tactics outside of SEO
-
Research & Trends
Dive into research and trends in the search industry.
-
Support
Connect on product support and feature requests.
Related Questions
-
Website blocked by Robots.txt in OSE
When viewing my client's website in OSE under the Top Pages tab, it shows that ALL pages are blocked by Robots.txt. This is extremely concerning because Google Webmaster Tools is showing me that all pages are indexed and OK. No crawl errors, no messages, no nothing. I did a "site:website.com" in Google and all of the pages of the website returned. Any thoughts? Where is OSE picking up this signal? I cannot find a blocked robots tag in the code or anything.
Moz Pro | | ConnellyPartners0 -
Moz Crawl Test: WordPress sites with and without /feed and /trackback entires?
I have multiple WP websites and on some of the websites, on my Moz Crawl test, I see an entry for every blog post but also entries for /feed and /trackback for that single blog post. For example, www...com/someArticle www....com/someArticle/feed www...com/someArticle/trackback 1. Can anyone explain why the Crawl test is picking up the /feed and /trackback items? Is it simply because they are 301 redirects to the original post (www...com/someArticle)? 2. What setting(s) in WordPress are making this information appear? Or is it just that the site(s) that have the /feed and /trackback are displaying "normal" behavior for a WP site with a lot of trackbacks and feed entires? 3. Should /fee and /trackback, as well as /author be blocked in robots.txt? Thanks in advance for your advice and input!
Moz Pro | | Titan5520 -
Since July 1, we've had a HUGE jump in errors on our weekly crawl. We don't think anything has changed on our website. Has MOZ changed something that would account for a large leap in duplicate content and duplicate title errors?
Our error report went from 1,900 to 18,000 in one swoop, starting right around the first of July. The errors are duplicate content and duplicate title, as if it does not see our 301 redirects. Any insights?
Moz Pro | | KristyFord0 -
Problem crawling a website with age verification page.
Hy every1, Need your help very urgent. I need to crawl a website that first has a page where you need to put your age for verification and after that you are redirected to the website. My problem is that SEOmoz, crawls only that first page, not the whole website. How can I crawl the whole website?, do you need me to upload a link to the website? Thank you very much Catalin
Moz Pro | | catalinmoraru0 -
Does it make sense to have multiple campaigns for one website?
I'm new to SEO and SEOMoz. While setting up my campaigns, I was a bit confused. Do I need to setup separate campaigns for different pages of my site? For example, I run a proofreading service: www.kibin.com I obviously want to track keyword on the main page (like proofreading service), however, I also want to track keywords (like 'essay editing') on www.kibin.com/essay-editing Should these be two separate campaigns or do I just put all the keywords into one campaign for kibin.com? Thanks!
Moz Pro | | Kibin0 -
Best Way to Include Social Media in Website?
Hi, how are you doing? I am new to MOZ, totally love it. I recently developed the social media for my page (www.aceromart.com) in Facebook, Google + and Twitter. (I am from Mexico, so the website is in Spanish) I am no expert in SEO whatsoever, but i like to engage my customers with great content both in my page and social media. My question is: **What the best way to include your social media links or icons on my page. Is there a program or a way to include the links. I want the people that visit ** **Should you include them in every page?, in a footer?, with icons or links. ** Thanks in advance for your advices, they are greatly appreciated. Best Regards, Jesus D
Moz Pro | | JesusD0 -
Strange Website Activity
Hello, I have been building websites for about 4 months now and finally had my first real success with a website. I found a niche that I was able to get on the first page with. This site was fine and then boom it dropped to #17 around the time of the recent Google changes. So I just thought it was that, I decided to add it as a campaign on here. This is when I noticed that the pages were not being crawled. So being new at all of this, I researched that. Of course, the main things were is the privacy and robots.txt. It is running on wordpress and I know not to set the blog to private, but I wasn't familiar with how to edit the robots.txt. I found a good plugin that easily allowed me to set the text to allow all bots. It seemed to be set to the normal wordpress settings before, and I never had a problem with any of my websites not being crawled. Anyway, once I just set it to allow all bots on ALL of my website, the pages started being crawled again. My traffic went back up and I was on the first page all day yesterday. Today, so far big drop off. So, I deleted the campaign and set up a new one. Sure enough no pages crawled yet. I made some security changes using Bulletproof Security and another plugin to see if that effects it. Nothing yet. I am just really confused as to what is going on, so if any of you have any ideas that would be great. It is a simple site, and I made some changes like theme when I was trying to figure out why the pages weren't being crawled. So it is not the most beautiful design right now. Also, I try my best to put up well-written useful content, so I don't think that is the issue for the rankings drops. I don't have many if any actual backlinks yet because of the newness of my site, could be the reason for it acting strange BUT none of that explains the pages not being crawled at the same time my site drops????? Sorry so long but had to explain it all! Thanks in advance to anyone who has anything to say about this situation! Edit: I should clarify, I put the the security plugins on yesterday while I was on the first page and have deleted them to see if that allows the pages to be crawled. Sorry if I wasn't clear.
Moz Pro | | iheartkelby0