Live data from Hacker News

Building the heap: racking 30 petabytes of hard drives for pretraining

si.inc

31–40 of 281 posts

Re: Building the heap: racking 30 petabytes of hard drives for pretraining

#31
post #27

$125/disk, 12k/mo depreciation cost which i assume means disk failures, so ~100 disks/mo or 1200/yr, which is half of their disks a year - seems like a lot.

no, we wanted to be conservative by depreciating somewhat more aggressively than that. we have much closer to 5% yearly disk failure rates.

Re: Building the heap: racking 30 petabytes of hard drives for pretraining

#32

Shows how crazy cheap on prem can be. tips hat

Not included is overhead of dealing with maintenance. S3/R2 generally don’t require OPS type dedicated to care and feeding. This type of setup will likely require someone to spend 5 hours a week dealing with it.

I once had about three racks full of servers under my control, admittedly they weren't a ton of disks, but still the hardware maintenance effort was pretty much negligible over a few years (until it all went to the cloud).

The majority of server wrangling work I spent dealing with OS updates and, most annoyingly, OpenStack. But that's something you can't escape even if you run your stuff in the cloud...

Re: Building the heap: racking 30 petabytes of hard drives for pretraining

#33
post #29
post #13

Earlier quoted context omitted.

True, though this is specifically for pretraining data (S3 wouldn't sell us used disk + no DR storage).

You're in a seismically active part of the world. Will the venture last in a total loss scenario?

We're currently 1/1 for the recent 4.3 magnitude earthquake (though if SF crumbles we might lose data)

Re: Building the heap: racking 30 petabytes of hard drives for pretraining

#34
post #27

$125/disk, 12k/mo depreciation cost which i assume means disk failures, so ~100 disks/mo or 1200/yr, which is half of their disks a year - seems like a lot.

It's an accounting term. You need to report the value of assets of your company each reporting cycle. This allows you to report company profit more accurately since the 2400 drives aren't likely not worth what the company originally paid. It's stated as a tax write-off but people get confused with that term (they think X written off == X less tax paid). It's better to correctly state it as a way to more accurately report profit (which may end up with less company tax paid but obviously not 1:1 since company tax is not 100%).

So anyway you basically pretend you resold the drives today. Here they are assuming in 3 years time no one will pay anything for the drives. Somewhat reasonable to be honest since the setup's bespoke and you'll only get a fraction of the value of 3 year old drives if you resold them.

Re: Building the heap: racking 30 petabytes of hard drives for pretraining

#36

Shows how crazy cheap on prem can be. tips hat

Not included is overhead of dealing with maintenance. S3/R2 generally don’t require OPS type dedicated to care and feeding. This type of setup will likely require someone to spend 5 hours a week dealing with it.

True, this is a large reason why we chose to have the datacenter a couple blocks away from the office.

Re: Building the heap: racking 30 petabytes of hard drives for pretraining

#37
Is it correct that you have zero data redundancy? This may work for you if you're just hoarding videos from YouTube, but not for most people who require an assurance that their data is safe. Even for you, it may hurt proper benchmarking, reproducibility, and multi-iteration training if the parent source disappears.

Re: Building the heap: racking 30 petabytes of hard drives for pretraining

#38

Is it correct that you have zero data redundancy? This may work for you if you're just hoarding videos from YouTube, but not for most people who require an assurance that their data is safe. Even for you, it may hurt proper benchmarking, reproducibility, and multi-iteration training if the parent source disappears.

Definitely much less redundancy, this was definitely a tradeoff we made for pretraining data and cost.

Re: Building the heap: racking 30 petabytes of hard drives for pretraining

#40

The biggest part that is always missing in such comparisons is the employee salaries. In the calculation they give $354k/year of total cost per year. But now add the cost of staff in SF to operate that thing.

someone has to go and power-cycle the machines every couple months it's chill, that's the point of not using ceph
Post reply on HN