Building the heap: racking 30 petabytes of hard drives for pretraining
91–100 of 281 posts
Re: Building the heap: racking 30 petabytes of hard drives for pretraining
#92Earlier quoted context omitted.
The disk failure rates are very low when compared to decade ago. I used to change more than a dozen disks every week a decade ago. Now it's an eyebrow raising event which I seldom see. I think following Backblaze's hard disk stats is enough at this point.
Backblaze reports an annual failure rate of 1.36% [0]. Since their cluster uses 2,400 drives, they would likely see ~32 failures a year (extra ~$4,000 annual capex, almost negligible). [0] https://www.backblaze.com/cloud-storage/resources/hard-drive...
2,400 drives. Mostly 12TB used enterprise drives (3/4 SATA, 1/4 SAS). The JBOD DS4246s work for either.
Re: Building the heap: racking 30 petabytes of hard drives for pretraining
#93IPMI is great and all, but I still prefer serial ports and remote PDUs. Never met a BMC I could trust.
Re: Building the heap: racking 30 petabytes of hard drives for pretraining
#94Earlier quoted context omitted.
The disk failure rates are very low when compared to decade ago. I used to change more than a dozen disks every week a decade ago. Now it's an eyebrow raising event which I seldom see. I think following Backblaze's hard disk stats is enough at this point.
Backblaze reports an annual failure rate of 1.36% [0]. Since their cluster uses 2,400 drives, they would likely see ~32 failures a year (extra ~$4,000 annual capex, almost negligible). [0] https://www.backblaze.com/cloud-storage/resources/hard-drive...
Re: Building the heap: racking 30 petabytes of hard drives for pretraining
#95Re: Building the heap: racking 30 petabytes of hard drives for pretraining
#96Re: Building the heap: racking 30 petabytes of hard drives for pretraining
#97'Networking was a substantial cost and required experimentation. We did not use DHCP as most enterprise switches don’t support it and we wanted public IPs for the nodes for convenient and performant access from our servers. While this is an area where we would have saved time with a cloud solution, we had our networking up within days and kinks ironed out within ~3 weeks.'
Where does the switch choice come into whether you DHCP? Wth would you want public IPs.
Re: Building the heap: racking 30 petabytes of hard drives for pretraining
#98Plus, the TCO is already way under the cloud equiv. so might as well spend a little more to get something much newer and more reliable
Re: Building the heap: racking 30 petabytes of hard drives for pretraining
#99i am still confused what their software stack is, they dont use ceph but bought netapp, so they use nfs?
Re: Building the heap: racking 30 petabytes of hard drives for pretraining
#100The networking stuff seems....odd. 'Networking was a substantial cost and required experimentation. We did not use DHCP as most enterprise switches don’t support it and we wanted public IPs for the nodes for convenient and performant access from our servers. While this is an area where we would have saved time with a cloud solution, we had our networking up within days and kinks ironed out within ~3 weeks.' Where doe…
So anyone can download 30 PB of data with ease of course.