Live data from Hacker News

10 petabytes - visualized

blog.backblaze.com

41–50 of 53 posts

Re: 10 petabytes - visualized

#41

Anyone want to guess how much data they actually have stored? Backblaze's pods are a data-loss nightmare -- lots of single points of failure which will wipe out many TB of data at a time -- and backblaze has stated that they replicate data across multiple pods. Given that the 10 PB seems to be the amount of raw storage backblaze has, I'm guessing that the amount of actual data stored is much less -- depending on what…

At SpiderOak we get a 3x replication equivalent for about 35% overhead, using Reed-Solomon at the cluster level (on top of RAID6 at the machine level.) Not nearly as expensive as outright replication. Agree those SATA port multiplies are worrisome. In the beginning, our prototype machines used them to squeeze as many drives into a single machine as possible. They have unusually low tolerance for electrical interferen…

Does SpiderOak only provide backup service? Erasure encoding is efficient for cold data. Do you use erasure encoding to distribute the hot data across clusters?

Re: 10 petabytes - visualized

#42
post #5

I am not american, nor do we use inches here in germany but when looking at #3 it should be noted that 5.75 inches is actually the drive length, not height as they state. Well depends on how you look at it but it confused me in the beginning.

I'm curious why this didn't get more upvotes. 5.75 is a diagonal I think?

Re: 10 petabytes - visualized

#43
post #28
post #19

Earlier quoted context omitted.

I wonder how many thousands of people are paying them to back up stuff they could easily redownload. Great business model if they de-dupe internally.

they encrypt all data and i dont think it would be wise of them analyzing their clients data for dupes...if somebody would get to know that their business would be gone.

ZFS supports encryption with deduplication, and Dropbox does dedupe accross customers (and why not? It's not a human digging through things).

http://blogs.sun.com/darren/entry/compress_encrypt_checksum_... - ZFS link

Re: 10 petabytes - visualized

#45

Earlier quoted context omitted.

At SpiderOak we get a 3x replication equivalent for about 35% overhead, using Reed-Solomon at the cluster level (on top of RAID6 at the machine level.) Not nearly as expensive as outright replication. Agree those SATA port multiplies are worrisome. In the beginning, our prototype machines used them to squeeze as many drives into a single machine as possible. They have unusually low tolerance for electrical interferen…

3x replication equivalent for about 35% overhead How did you compute this "replication equivalent"?

Picked a number from thin air? Raid six requires 2 drives for back and is normally used in set's of 8 or 16 drives but looks like they are using 45 drives. So 45/43 = 4.65% overhead from using RAID.

Not if they lose 35% on top of that they are around 41% overhead. But, they are taking a huge it on write speeds, network traffic and reliability for doing so.

Edit: Looks like they have 10,058 TB before partitioning the drives so my guess is ~3-6TB of actual user data.

Re: 10 petabytes - visualized

#46
post #3

Is there any reason why 10 PB was split in that particular combination of 2, 1.5 and 1TB drives? EDIT: And, just a comparison: Google processes over 20PB of data per day!

and why it takes more 2TB drives than 1TB drives...

That's the drives they actually have.

2280 * 2 + 3166 * 1.5 + 1 * 749 = 10,058 TB.

Re: 10 petabytes - visualized

#47
I had no idea Backblaze had gotten this popular, but I'm glad to see their business growing. If you happen to have the profile they've optimized for, it really is the best thing going.

Re: 10 petabytes - visualized

#48
post #6

I'll give you a much more compact way to envision 10 Pb: 10 Petabytes is 10,000,000,000,000,000 bytes or 80,000,000,000,000,000 bits, divided by 8,000,000,000 bits per full human genome (2 bits per base-pair) that's about 10,000,000 cells (not red blood cells because they don't contain DNA), or about 5 milliliters! (10 um diameter on average so about 500 cubic um, so 2 million or so per ml), and that includes all the…

  -1 already eh? [..]
You really oughta know that you should wait a bit for stuff to get overaged over a few hours.

Re: 10 petabytes - visualized

#49
post #28

Earlier quoted context omitted.

they encrypt all data and i dont think it would be wise of them analyzing their clients data for dupes...if somebody would get to know that their business would be gone.

ZFS supports encryption with deduplication, and Dropbox does dedupe accross customers (and why not? It's not a human digging through things). http://blogs.sun.com/darren/entry/compress_encrypt_checksum_... - ZFS link

When different customers use different key to encrypt the data, even if the source is the same, the outcome is different. So it is not feasible to do de-duplication across different customers' data, unless all of them use the same key.

Re: 10 petabytes - visualized

#50
post #40
post #30

Amazing, a discussion of data without comparing the internationally accepted unit of "library of congress". I'm more interested in how many LoCs 10 Pb is? If we built a city of LoCs, how many acres would that city be? These are the questions that keep me up late at night.

The Library of Congress is 10TB. Zipped, it would fit on a single 3TB hard drive.

Perhaps that implies we should start using kiloLoC and megaLoCs?
Post reply on HN