Live data from Hacker News

Free 5 Billion Page Web Index Now Available from Common Crawl Foundation

readwriteweb.com

1–10 of 42 posts

Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation

#2
One interesting discussion from here: http://www.commoncrawl.org/common-crawl-enters-a-new-phase/ It says the cost of running a hadoop job to scan all 5billon documents is in the order of $100.

Does any one know how does this compare to let say Yahoo BOSS? Is it even comparable?

Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation

#3
post #2

One interesting discussion from here: http://www.commoncrawl.org/common-crawl-enters-a-new-phase/ It says the cost of running a hadoop job to scan all 5billon documents is in the order of $100. Does any one know how does this compare to let say Yahoo BOSS? Is it even comparable?

Does BOSS still exist? I was under the impression that it was defunct.

Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation

#4
I wonder if crooks will try to exploit this crawl. As a person who has an index of the web like this it has been interesting to see what they look for. SSN's and credit card numbers are common, as are sites running older versions of PHP software or exploitable shopping carts.

Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation

#5
post #3
post #2

One interesting discussion from here: http://www.commoncrawl.org/common-crawl-enters-a-new-phase/ It says the cost of running a hadoop job to scan all 5billon documents is in the order of $100. Does any one know how does this compare to let say Yahoo BOSS? Is it even comparable?

Does BOSS still exist? I was under the impression that it was defunct.

Yes. With Google no longer providing search result API (not even paid version, the last I checked) people are turning to BOSS/Bing/(anything else?)

Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation

#8
post #4

I wonder if crooks will try to exploit this crawl. As a person who has an index of the web like this it has been interesting to see what they look for. SSN's and credit card numbers are common, as are sites running older versions of PHP software or exploitable shopping carts.

It makes it very easy for people to steal vast amounts of your content and republish it on their own sites, with ads all around it.

Many content sites have protections in place to recognize bots by their behavior or use "honeypots" to tell bots apart from human visitors and thus avoid large scale content theft.

Re: Free 5 Billion Page Web Index Now Available from Common Crawl Foundation

#9
post #7

I'd love to see Gabriel weigh in on this. I wonder if Duck Duck Go will be able to take advantage of this resource?

Nova Spivack said that the crawls have been going for several years. There's a good chance that many of the pages in the archive are unacceptably outdated for indexing purposes.
Post reply on HN