Shows how crazy cheap on prem can be. tips hat
Not included is overhead of dealing with maintenance. S3/R2 generally don’t require OPS type dedicated to care and feeding. This type of setup will likely require someone to spend 5 hours a week dealing with it.
Building the heap: racking 30 petabytes of hard drives for pretraining
41–50 of 281 posts
Re: Building the heap: racking 30 petabytes of hard drives for pretraining
#42Earlier quoted context omitted.
True, though this is specifically for pretraining data (S3 wouldn't sell us used disk + no DR storage).
You're in a seismically active part of the world. Will the venture last in a total loss scenario?
Re: Building the heap: racking 30 petabytes of hard drives for pretraining
#43I love this story. This is true hacking and startup cost awareness.
Re: Building the heap: racking 30 petabytes of hard drives for pretraining
#44No mention of disk failure rates? curious how it's holding up after a few months
The disk failure rates are very low when compared to decade ago. I used to change more than a dozen disks every week a decade ago. Now it's an eyebrow raising event which I seldom see. I think following Backblaze's hard disk stats is enough at this point.
[0] https://www.backblaze.com/cloud-storage/resources/hard-drive...
Re: Building the heap: racking 30 petabytes of hard drives for pretraining
#45Re: Building the heap: racking 30 petabytes of hard drives for pretraining
#46Re: Building the heap: racking 30 petabytes of hard drives for pretraining
#47Earlier quoted context omitted.
Not included is overhead of dealing with maintenance. S3/R2 generally don’t require OPS type dedicated to care and feeding. This type of setup will likely require someone to spend 5 hours a week dealing with it.
I once had about three racks full of servers under my control, admittedly they weren't a ton of disks, but still the hardware maintenance effort was pretty much negligible over a few years (until it all went to the cloud). The majority of server wrangling work I spent dealing with OS updates and, most annoyingly, OpenStack. But that's something you can't escape even if you run your stuff in the cloud...
$LastJob we ran a ton of Azure Web App Containers, alot of OS work no longer existed so it's possible with Cloud to remove alot of OS toil.
Re: Building the heap: racking 30 petabytes of hard drives for pretraining
#48The biggest part that is always missing in such comparisons is the employee salaries. In the calculation they give $354k/year of total cost per year. But now add the cost of staff in SF to operate that thing.
someone has to go and power-cycle the machines every couple months it's chill, that's the point of not using ceph
Re: Building the heap: racking 30 petabytes of hard drives for pretraining
#49I'm curious about the process of getting colo space. Did you use a broker? Did you negotiate, and if so, how large was the difference in price between what you initially were quoted and what you ended up paying?
Re: Building the heap: racking 30 petabytes of hard drives for pretraining
#50The biggest part that is always missing in such comparisons is the employee salaries. In the calculation they give $354k/year of total cost per year. But now add the cost of staff in SF to operate that thing.
someone has to go and power-cycle the machines every couple months it's chill, that's the point of not using ceph
I think that anyone with actual experience of operating thousands of physical disks in datacenters would challenge this assumption.