Just wanted to say, thanks for doing this! Now the old rant... I started my career when on-prem was the norm and remember so much trouble. When you have long-lived hardware, eventually, no matter how hard you try, you just start to treat it as a pet and state naturally accumulates. Then, as the hardware starts to be not good enough, you need to upgrade. There's an internal team that presents the "commodity" interface…
Building the heap: racking 30 petabytes of hard drives for pretraining
81–90 of 281 posts
Re: Building the heap: racking 30 petabytes of hard drives for pretraining
#82Just wanted to say, thanks for doing this! Now the old rant... I started my career when on-prem was the norm and remember so much trouble. When you have long-lived hardware, eventually, no matter how hard you try, you just start to treat it as a pet and state naturally accumulates. Then, as the hardware starts to be not good enough, you need to upgrade. There's an internal team that presents the "commodity" interface…
Re: Building the heap: racking 30 petabytes of hard drives for pretraining
#83Earlier quoted context omitted.
They mentioned the cluster being used enterprise drives, I can see the desire to save money but agree, that is going to be one expensive mistake down the road. I should also note personally for home cluster use, I learned quickly that used drives didn’t seem to make sense. Too much performance variability.
If I remember correctly, most drives either: 1. Fail in the first X amount of time 2. Fail towards the end of their rated lifespan So buying used drives doesn't seem like the worst idea to me. You've already filtered out the drivers that would fail early. Disclaimer: I have no idea what I'm talking about
Re: Building the heap: racking 30 petabytes of hard drives for pretraining
#84Earlier quoted context omitted.
> The amount of time lost to driving to the datacenter, waiting for replacement parts to arrive, and scrambling to patch over unexpected failure modes is always much higher than expected. I don't have this experience at all. Our colo handled almost all work. the only time i ever went to the server farm was to build out whole new racks. Even replacing servers the colo handled for us at good cost. Our reliability came…
> Our reliability came from software not hardware, though of course we had hundreds of spares sitting by, the defense in depth (multiple datacenters, each datacenter having 2 'brains' which could hotswap, each client multiply backed up on 3-4 machines)... Of course, but building and managing the software stack, managing hundreds of spares across locations, spanning across datacenters, having a hotswap backup system i…
Re: Building the heap: racking 30 petabytes of hard drives for pretraining
#85Earlier quoted context omitted.
They can rent a dark fiber for themselves for that distance, and it'll be cheap. However, as they noted they use 100gbps capacity from their ISP.
Does San Francisco really still have dark fiber? That 90s bubble sure did overshoot demand.
Re: Building the heap: racking 30 petabytes of hard drives for pretraining
#86Earlier quoted context omitted.
we have 6 months of experience operating thousands of physical disks in datacenters now! it's about a couple hours a month of employee time in steady-state.
How about all the other infrastructure. Since you are obviously not using the cloud, you must have massive amounts of GPUs and operating systems. All of that has been working together, it's not just keep watching for the physical disks and all is set. Don't get me wrong, I buy the actual numbers regarding hardware costs, but in addition to that presenting the rest as basically a one man show in terms of maintenance h…
we also use cloudflare extensively for everything that isn't the core heap dataset, the convenience of buckets is totally worth it for most day-to-day usage.
the heap is really just the main pretraining corpus and nothing else.
Re: Building the heap: racking 30 petabytes of hard drives for pretraining
#87Earlier quoted context omitted.
Definitely much less redundancy, this was definitely a tradeoff we made for pretraining data and cost.
Did you do any kind of redundancy at least (eg: putting every 10 disks in RAID 5 or RAID Z1)? Or I suppose your training application doesn't mind if you shed a few terabytes of data every so often?
Re: Building the heap: racking 30 petabytes of hard drives for pretraining
#88No mention of disk failure rates? curious how it's holding up after a few months
Re: Building the heap: racking 30 petabytes of hard drives for pretraining
#89Earlier quoted context omitted.
True, though this is specifically for pretraining data (S3 wouldn't sell us used disk + no DR storage).
I do appreciate the scrappiness of your solution. Used drives for a storage cluster is like /r/homelab on steroids. And since it's pretraining data, I suppose data integrity isn't critical. Most venture-backed startups would have just paid the AWS or Cloudflare tax. I certainly hope your VCs appreciate how efficient you are being with their capital :)
Re: Building the heap: racking 30 petabytes of hard drives for pretraining
#90Earlier quoted context omitted.
They can rent a dark fiber for themselves for that distance, and it'll be cheap. However, as they noted they use 100gbps capacity from their ISP.
We want to get darkfiber from the datacenter to the office. I love 100Gbps