Java or C++
Our new search index: Caffeine
21–30 of 82 posts
Re: Our new search index: Caffeine
#22Re: Our new search index: Caffeine
#23Nice to see they reference an iPod instead of a Nexus One. Somehow I feel that if it was a Microsoft announcement, the company would reference Zunes.
The Nexus One is a phone, not a device with a large storage capacity. If Google made one of those, I'm sure it would be referenced.
edit: the Nexus One box also includes a 4GB micro SD card, but that'd be comparing flash memory with harddrives, and is hardly an intuitive explanation.
Re: Our new search index: Caffeine
#24That is an enormous amount of data. I wonder how much is junk?
You're new to the Internetz, right? :)
Re: Our new search index: Caffeine
#25+5 for an intuitive analogy. Priceless!
Re: Our new search index: Caffeine
#26new index: solution: caffeine, Google's new index will analyze the web in small portions and update its search index on a continuous basis, globally. As Google find new pages, or new information on existing pages, Google can add these straight to the index. improvement: user can find fresher information than ever before—no matter when or where it was published.
recent changes in Google search result page suggested that the searched results appears to be categorized further to the different structured data. This could be inline with the rolling out Caffeine.
Re: Our new search index: Caffeine
#27100 petabytes of information (100 million GiB) feeding their index is more than I would have expected.
On the other hand it doesn't seem that much when you consider todays storage density. You can fit around 0.5 PB into one rack nowadays. 200 racks then sounds a bit less impressive than 100 Petabytes. However, that ofcourse doesn't account for redundancy, nor for doing anything useful with such a pile of data. Both of which impose some interesting challenges at that scale.
It is to laugh. I don't think that's been true since around the time they stopped building server shelves out of lego or cardboard.
They've been keeping the full text of the web in RAM (for the snippets), with indexes, several times over. With independent live siblings in multiple datacenters.
Re: Our new search index: Caffeine
#28Wow is that ever a terrible infographic.
Re: Our new search index: Caffeine
#29Earlier quoted context omitted.
On the other hand it doesn't seem that much when you consider todays storage density. You can fit around 0.5 PB into one rack nowadays. 200 racks then sounds a bit less impressive than 100 Petabytes. However, that ofcourse doesn't account for redundancy, nor for doing anything useful with such a pile of data. Both of which impose some interesting challenges at that scale.
You think Google's mainline storage of their index is on disk ? It is to laugh. I don't think that's been true since around the time they stopped building server shelves out of lego or cardboard. They've been keeping the full text of the web in RAM (for the snippets), with indexes, several times over. With independent live siblings in multiple datacenters.
I merely tried to put the figure into perspective.
Re: Our new search index: Caffeine
#30Earlier quoted context omitted.
On the other hand it doesn't seem that much when you consider todays storage density. You can fit around 0.5 PB into one rack nowadays. 200 racks then sounds a bit less impressive than 100 Petabytes. However, that ofcourse doesn't account for redundancy, nor for doing anything useful with such a pile of data. Both of which impose some interesting challenges at that scale.
You think Google's mainline storage of their index is on disk ? It is to laugh. I don't think that's been true since around the time they stopped building server shelves out of lego or cardboard. They've been keeping the full text of the web in RAM (for the snippets), with indexes, several times over. With independent live siblings in multiple datacenters.
Let's suppose for a second that they can get 8GB sticks of RAM for 1/3 the retail price of around $600, and for the sake of the back of the envelope calculation, let's round it up to ten GB. So let's call it $20/GB of RAM. Four gigabyte sticks might seem cheaper until you consider they'd have to double the number of machines to hold them. Now they mentioned that the size of the index is 100,000,000 GB, which means that they would have spent $2 billion on this RAM alone, not to mention all the other components that are required to house it.
Especially considering how much their capital expenditure has fallen lately (http://www.datacenterknowledge.com/archives/2010/01/22/googl...), it doesn't seem very likely that they could afford to spend even $2 billion on this. And the assumptions I made about the price of RAM are pretty generous, considering they'd have had to have been acquiring this for some time, and therefore that the RAM would have been more expensive previously.