Live data from Hacker News

Ask HN: How would you store 10PB of data for your startup today?

news.ycombinator.com

301–310 of 374 posts

Re: Ask HN: How would you store 10PB of data for your startup today?

#301
post #277
post #264

Earlier quoted context omitted.

> 625 16TB drives is $400000 how much is the real estate cost of 625 drives and associated machinery to run it? At a guess, AWS has an operating margin of about 30%, so you can approximate their cost of hardware, bandwidth, and other fixed costs as 70% of their sticker price. As a start up, can you actually get this price to be lower? I actually dont think you can, unless your operation is very small and can be done…

Their margin on bandwidth is literally over 1000%. Quick google says that S3 costs 320% more than Backblaze (which, presumably, isn't running at a loss). > At a guess The comments in this discussion that try to provide actual numbers show a fairly lopsided argument against S3. The comments that are advocating for S3 aren't as detailed. You can look at this at the macro level, as on comment did, and see that one 1.2PB…

> one 1.2PB RackmountPro 4U server is $11K

An empty 4U server with 96+ bays looks like it will set you back ~$7k minimum. At $500 per drive (I have no idea what volume discounts are like) filling it with drives would be in the range of ~$50k. You'd still need RAM. And (as you noted) space and power.

I have no idea how the math ends up working out, but a 1PB appliance in working order is nowhere near as cheap as $11k.

Re: Ask HN: How would you store 10PB of data for your startup today?

#302

Earlier quoted context omitted.

GPT-3 that convinces you why you don't need to store this thing

Image processing to label each image ("baby with spaghetti on head", "cat playing with string", "naked person"), and then only save one image with each label.

Just save the label. Then you use generative techniques when they want to retrieve the image.

Re: Ask HN: How would you store 10PB of data for your startup today?

#303
I'd like to echo an suggestion I read earlier in this thread: at this scale (i.e. yearly spent), talk to AWS, GCP, Azure or a reseller of your trust and get a good deal to compare your other options with.

Disclaimer: I'm working at a consultancy/partner for a competing cloud.

Re: Ask HN: How would you store 10PB of data for your startup today?

#304

Earlier quoted context omitted.

Totally out-of-band for this thread, but... what are the uses for a multi-gigabyte key?! I'm clearly unaware of some cool tech, any key words I can search?

When I say "key" I mean the blob that gets stored, but I may be misremembering or misusing S3 terms... it was large amounts of DNA sequencing data, and one of the first tasks was to add S3 support for indexed reads to our internal HTSlib fork, and since then somebody else's implementation has been added to the library. In any case, I quickly forgot about most of the details of S3 when I no longer had to deal with it…

That makes perfect sense, assume I'm a pleb! Thanks for the follow-up, large files/values make sense.

My head was in huge cryptographic keys for some purpose

Re: Ask HN: How would you store 10PB of data for your startup today?

#305
Need more details.. maybe a graph (or several graphs ) of requests \ day for various items (categorized by popularity and size is ok ) (a curve ( i suppose not very hyperbolic) to breakdown populary of top requested items vs long tail of almost never seen, and rarely seen which i suppose is the most 9f those 10pb ) and current bandwidth intersection ( and data size ) and volume , this is too have an idea about bw, iops ,structure of the data and requests patterns and requirements and caching layer , i think that probably a share fs is worse than distributed blobl storage here ( assuming spinning disks somewhere and not huge caches ) Not all days usage patterns are equal, your requirement are different from database (which is more in line with some suggestions here ) Plus data safety is everything for your kind of business so redoundancy is a must , speed too (don't even think about filecoin imho) i would think more about a mix of spinning and name as cache layer redoundant on multiple datacenter if it's to save costs.. if it's to save efforts and a bit of costs look at ovh offerings for blob storage services or contact backblaze for a custom solution hosted by them ?

Plus here we are not talking about 10pb but probably at 25 given redoundancy and probably also at 100pb ad more given the assumption that your company is growing , so a solution that cost slightly less today but will only do 2x when you do 10x would still be very interesting imo.. there is a lot to talk about ;)

Re: Ask HN: How would you store 10PB of data for your startup today?

#306
I have done 2 PB HPC data storage with ZFS. If I may extrapolate, I don’t see why it wouldn’t workout the same for 10 PB.

A 1U rack server attached to two JBODs(each 4U containing 60 spinning disks) connected to the server via 4 SAS HD cables. The rack server gets 512GiB of RAM to cache reads, and an Optane drive as persistent cache for writes. The usable storage depends on your redundancy and spare needs. But, as an example my setup - (9 * 6 drives(RAIDz2) + 4 hot spares) nets me about 450 TiB per JBOD or 900 TiB per rack server + two JBODs.

Repeat the setup by 6 times, and it would meet your 10 PB need. Throw in a few links 10GBps per server and have them all linked up by a switch, and you got your own storage setup. May be Minio(I have no experience with it) or something like that would give you a S3 interface over the whole thing.

I bet it would come out much cheaper than AWS. But, you’ve got to get your hands dirty a bit with system in work, and automate all the things with a tool like Ansible. Having done it, I’d say it is totally worth it at your scale.

Re: Ask HN: How would you store 10PB of data for your startup today?

#307
At startup grade, it‘s fine to stick and grow with IaaS provider like Amazon, Google, Microsoft, Oracle or whatever you like.

However, you‘ll get to a point, where it‘s crucial to become profitable. And storing that much data does cost a lot of money using one of the mentioned providers.

So, when you think it‘s the right time to become “mature”, then get your own servers up and running using colocation.

What options do you have here (just a quick brainstorm): 1. Set up some servers, put in a lot of hard drives, format them using zfs and make it available using nfs on your network 2. Get some storage servers 3. Set up a Ceph cluster

I used to work as a CTO at a hosting company and evaluated all of these options and more. Every of these options comes with pros and cons.

Just one last advice: Evaluate your options and get some external help on this. Any of these options have pitfalls and you need experienced consultants to set up and run such an infrastructure.

All in all, it’s an invest, that will save you a lot of money and will give you freedom and flexibility to grow further.

P.S. we ended up setting up a Ceph cluster. We found a partner, who’s specialized on hosting custom infrastructures. That partner is responsible for all the maintenance, so we could focus on the product itself.

Re: Ask HN: How would you store 10PB of data for your startup today?

#308
post #262
post #212

I don't really know much about optimizing storage costs, But You could learn from storage giants. Example is Blackblaze storage pod 6.0 according to them it holds 0.5PB with a cost of 10k$, you will need about 20*10K$ = 200K$ + Maintenance(They also publish failure rates) , The schematics and everything is in their website and according to them they have already a supplier who provides them with such devices which yo…

fyi that 10k includes no drives.

Reading: https://www.backblaze.com/blog/open-source-data-storage-serv... it seems the drives are included.

That 10.3k includes drives but you have to assemble the pod yourself.

For 12.8k you get drives and assembled pod from 3rd party manufacturer.

Backblaze pays about 8.7k at a scale for the whole enchilada

Those numbers do not make sense if we exclude drives. The server itself is not that expensive(2-3k tops) without the drives.

Re: Ask HN: How would you store 10PB of data for your startup today?

#309
post #263

Backblaze B2, ingress and egress are free through Cloudflare, and it's S3 compatible. It's peanuts by comparison but I've been storing ~22TB on there for years and love it. Wasabi and Glacier would be my 2nd choices.

Wait what?! There is a way to egress from free from Backblaze B2! That’s a big deal if true.

Yup, B2 egress has been free through CloudFlare for years:

- https://www.backblaze.com/blog/backblaze-and-cloudflare-part...

- https://www.cloudflare.com/partners/technology-partners/back...

- https://www.cloudflare.com/bandwidth-alliance/backblaze/

Re: Ask HN: How would you store 10PB of data for your startup today?

#310
post #90

Non-cloud: HPE sells their Apollo 4000[^1] line, which takes 60x3.5" drives - with 16TB drives, that's 960TB each machine, one rack of 10 of these is 9PB+ therefore, which nearly covers your 10PB needs. (We have some racks like this). They are not cheap. (Note: Quanta makes servers that can take 108x3.5" drive, but they need special deep racks.) The problem here would be the "filesystem" (read: the distributed servic…

When we set up user content storage of images and mp3s for Last.fm back in 2006ish we used MogileFS (from the bradfitz LJ perl days) running on our own hardware. 3/4/5/6u machines stuffed full of disks. I still think it's an elegant concept – easy to grok, easy to debug, easy to reason about. No special distributed filesystem to worry about.

Don't take this as an endorsement of the MogileFS perl codebase in 2021, but worth considering this style of storage system depending on your precise needs.

Post reply on HN