Live data from Hacker News

How would you build an internet scale web crawler?

news.ycombinator.com

11–16 of 16 posts

Re: How would you build an internet scale web crawler?

#11
post #8

I worked at Alexa, then affiliatated with the Internet Archive, on exactly this, in the late 90s. We built our own server farm to crawl, process and store the data. We had 30TB of storage, holding three "snapshots" of the web, and thought we were pretty hot stuff. That would sit on your desktop today. Crawling was the easy part. We had two processes of up to 40 threads each bringing the data down. Even this we had to…

impressive thanks for the wonderful read.

Re: How would you build an internet scale web crawler?

#14
post #8

I worked at Alexa, then affiliatated with the Internet Archive, on exactly this, in the late 90s. We built our own server farm to crawl, process and store the data. We had 30TB of storage, holding three "snapshots" of the web, and thought we were pretty hot stuff. That would sit on your desktop today. Crawling was the easy part. We had two processes of up to 40 threads each bringing the data down. Even this we had to…

“Alexa could have been Google, but went in another direction.“

Curious, why did they go in another direction?

Re: How would you build an internet scale web crawler?

#15
Concerning the crawling aspect of this, Michael Nielsen wrote a great article about how to download/crawl a significant part of the Internet in 40 hours:

http://www.michaelnielsen.org/ddi/how-to-crawl-a-quarter-bil...

It's a bit dated (from 2012) but probably still relevant.

Re: How would you build an internet scale web crawler?

#16
post #14
post #8

I worked at Alexa, then affiliatated with the Internet Archive, on exactly this, in the late 90s. We built our own server farm to crawl, process and store the data. We had 30TB of storage, holding three "snapshots" of the web, and thought we were pretty hot stuff. That would sit on your desktop today. Crawling was the easy part. We had two processes of up to 40 threads each bringing the data down. Even this we had to…

“Alexa could have been Google, but went in another direction.“ Curious, why did they go in another direction?

Amazon's stated reason for acquiring Alexa was to use Alexa's technology to build a recommendation engine. Search was never a priority for Alexa itself. We are acquired for $100 million, so it was take the money and run.
Post reply on HN