I worked at Alexa, then affiliatated with the Internet Archive, on exactly this, in the late 90s. We built our own server farm to crawl, process and store the data. We had 30TB of storage, holding three "snapshots" of the web, and thought we were pretty hot stuff. That would sit on your desktop today. Crawling was the easy part. We had two processes of up to 40 threads each bringing the data down. Even this we had to…
How would you build an internet scale web crawler?
11–16 of 16 posts
Re: How would you build an internet scale web crawler?
#12Re: How would you build an internet scale web crawler?
#13Re: How would you build an internet scale web crawler?
#14I worked at Alexa, then affiliatated with the Internet Archive, on exactly this, in the late 90s. We built our own server farm to crawl, process and store the data. We had 30TB of storage, holding three "snapshots" of the web, and thought we were pretty hot stuff. That would sit on your desktop today. Crawling was the easy part. We had two processes of up to 40 threads each bringing the data down. Even this we had to…
Curious, why did they go in another direction?
Re: How would you build an internet scale web crawler?
#15http://www.michaelnielsen.org/ddi/how-to-crawl-a-quarter-bil...
It's a bit dated (from 2012) but probably still relevant.
Re: How would you build an internet scale web crawler?
#16I worked at Alexa, then affiliatated with the Internet Archive, on exactly this, in the late 90s. We built our own server farm to crawl, process and store the data. We had 30TB of storage, holding three "snapshots" of the web, and thought we were pretty hot stuff. That would sit on your desktop today. Crawling was the easy part. We had two processes of up to 40 threads each bringing the data down. Even this we had to…
“Alexa could have been Google, but went in another direction.“ Curious, why did they go in another direction?