Live data from Hacker News

10 petabytes - visualized

blog.backblaze.com

51–53 of 53 posts

Re: 10 petabytes - visualized

#51
post #39

Earlier quoted context omitted.

- There is no way to retrieve information in this system. DNA is natively a content addressable storage system, due to natural base-pairing. But to first address your question in your edit: think of DNA as a pair of singly-linked lists, each with an alphabet of four characters. Each singly-linked list is the "reverse complement" of the other: the reverse sequence with an A-T swap and a C-G swap. At each position ther…

> To probe for information in a DNA database, you construct the reverse-complement of the desired bit of information To build the reverse-complement of the information, don't you need to have the information in the first place? I know it's possible to store information in DNA through various means, but I don't believe it can be done at the density the OP calculated. If we're going to take into account only informatio…

Think of it as a massive key-value store: you construct your query off the key, use that to pull out the key-value pair, and when you sequence your key you continue to sequence more in order to pull out the value. If you prefer sequential addresses, your key could be just that.

And actually, this could be done at a much higher density than what the original poster described, as he's counting the full cell in the density calculation, and DNA is only a small fraction of the cellular volume. You could duplicate all the DNA 10-100 times in the same amount of space once you take out all the ribosomes, proteins and extra water. And as long as it's not stored in direct sunlight or next to your pile of plutonium, DNA is going to be much much more stable than aligning magnetic fields. We're still getting good DNA sequence out of bones that are tens of thousands of years old.

When you think of nanotechnology and miniaturization, think of biology, because that's where all the real nanotechnology is going on. We've not done any better than nature when it comes to making small machinery. Nature has already invented the commodity interchangeable parts (amino acids and nucleic acids) that can self-assemble into rather fantastic machines.

However, we have beaten mother nature on latency: as I alluded to, a DNA database like this would have latency on the order of days for a lookup. On the other hand, as much parallel access as you can imagine is built in, without additional volume. And this isn't a system that has been engineered at all, I'm just talking about the fundamental properties of a little puddle of DNA and water. If half the engineering that went into modern computer hardware were put into a DNA database, it could be quite competitive with our electronic systems.

Re: 10 petabytes - visualized

#52
post #39

Earlier quoted context omitted.

> To probe for information in a DNA database, you construct the reverse-complement of the desired bit of information To build the reverse-complement of the information, don't you need to have the information in the first place? I know it's possible to store information in DNA through various means, but I don't believe it can be done at the density the OP calculated. If we're going to take into account only informatio…

Think of it as a massive key-value store: you construct your query off the key, use that to pull out the key-value pair, and when you sequence your key you continue to sequence more in order to pull out the value. If you prefer sequential addresses, your key could be just that. And actually, this could be done at a much higher density than what the original poster described, as he's counting the full cell in the dens…

Sure, a key-value store would work. My point is that the OP's system is not such a store. He just stores 10 PBs of raw data with no indexing and no duplication, so there is no way to retrieve data and comparison with hard disks is meaningless. My post was an answer to his "please correct my math".

Re: 10 petabytes - visualized

#53

My 350GB of music is backed up on Backblaze (a wonderful service), but seeing those visualizations reminded me of the environmental impact of my digital packratting. That's a sizable data center keeping my data happy and backed up.

Off topic, but what do you listen to that takes up 350GB? Is everything in FLAC?

I listen to a lot of music (from all genres). I only purchase/download full albums (even if most of the songs are not very good), and only listen to at least 196k MP3 (most are at least 256k or higher). Some are Apple Lossless (converted from FLAC). And throw in a few discographies in there and all of a sudden you have 350GB. I've discovered that I add about 50GB/6 months.
Post reply on HN