Live data from Hacker News

Building the heap: racking 30 petabytes of hard drives for pretraining

si.inc

251–260 of 281 posts

Re: Building the heap: racking 30 petabytes of hard drives for pretraining

#251
post #40

Earlier quoted context omitted.

someone has to go and power-cycle the machines every couple months it's chill, that's the point of not using ceph

You are under the assumption that only Ceph (and similar complex software) requires staff, whereas plain 30 PB can be operated basically just by rebooting from time to time. I think that anyone with actual experience of operating thousands of physical disks in datacenters would challenge this assumption.

Not really. Have spare drives on the shelf and use the "remote-hands" feature from the CoLo provider. Just open a ticket to have the drive swapped. Pretty easy. For remote server connections just use IPMI/iKVM and iPXE. Again, not too difficult.

The biggest hurdle is getting a mgmt system in place to alert you when something goes wrong - especially at this size. Grafana, Loki, monit, etc are all good tools to leverage that provide quick fault identification.

Re: Building the heap: racking 30 petabytes of hard drives for pretraining

#252
post #45

It's quite cheap to just store data at rest, but I'm pretty confused by the training and networking set up here. It sounds like from other comments that you're not going to put the GPUs in the same location, so you'll be doing all training over X 100 Gbps lines between sites? Aren't you going to end up totally bottlenecked during pretraining here?

30PB / 100Gbps comes down to about a month, 4 links would give you a week, so that seems pretty quite acceptable for a training run, especially since you can overlap the initial loading of the array with the first training, i.e train as data becomes available.

It goes without saying any data pre-processing needs to be done before writing, at the storage site, or on the training GPUs.

Re: Building the heap: racking 30 petabytes of hard drives for pretraining

#253
post #48

Earlier quoted context omitted.

So the drives are never going to fail? PSUs are never going to burn out? You are never going to need to procure new parts? Negotiate with vendors?

This concern troll that everyone trots out when anyone brings up running their own gear is just exhausting. The hyperscalers have melted people’s brains to a point where they can’t even fathom running shit for themselves. Yes, drives are going to fail. Yes, power supplies are going to burn out. Yes, god, you’re going to get new parts. Yes, you will have to actually talk to vendors. Big. Deal. This shit is -not- hard.…

Thanks for this. I agree, there seems to be some sort of resistance to building and maintaining a CoLo infrastructure. In reality, it is not too difficult. As I mentioned above, spare parts on the shelve with the CoLo "remote hands" support and a good monitoring system can lessen the impact of almost any catastrophic issue.

For the record, I have built (and currently maintain) a number of CoLo deployments. Our systems have been running for +10 years with very little failure of either drives or PSUs. In fact, our PSU failure rate is probably 1 every 3-4 years, and we probably loose a couple of drives per year. All in all, the systems are very reliable.

Re: Building the heap: racking 30 petabytes of hard drives for pretraining

#254
post #131

For a workload of that size you would be able to negotiate private pricing with AWS or any cloud provider, not just CloudFlare. You can get a private pricing deal on S3 with as little as half a PB. Not saying that your overall expenses would be cheaper w/a CSP than DIY, but its not exactly an apples to apples comparison of taking full retail prices for the CSPs against eBayed equipment and free labor (minus the cost…

What sort of deal are you taking about? Would it be 50% or more?

You can get way higher than 50% discounts with AWS (or any cloud) depending upon the scale of the buy.

Re: Building the heap: racking 30 petabytes of hard drives for pretraining

#255

Where does one acquire 90M hours of video without being YouTube?

My guess is automated surveillance, which is also where this whole play has to be headed.

Seems like that would be a good niche, not only for avoiding massive copyright considerations.

Also, it's some of the most boring footage where there's overwhelming amounts that's about the least desirable thing for humans to sit and watch every minute of.

Why send a human to do a machine's job?

Re: Building the heap: racking 30 petabytes of hard drives for pretraining

#256
post #162

Earlier quoted context omitted.

I’m also curious about this. I don’t recall seeing that mentioned in the article

Its in the first sentence: "We built a storage cluster in downtown SF to store 90 million hours worth of video data."

They were asking for the source of those data

Re: Building the heap: racking 30 petabytes of hard drives for pretraining

#257
post #250

Earlier quoted context omitted.

Small startup teams can sometimes get away with datacenter management being a side task that gets done on an as-needed basis at first. It will come with downtime and your stability won't be anywhere near as good as Cloudflare or AWS no matter how well you plan, though. Every real-world colocation or self-hosting project I've ever been around has underestimate their downtime and rate of problems by at least an order o…

For drive issues, this is easy. Have a stack of replacements on hand and just open a "remote-hands" ticket with the CoLo provider to swap out the drive. This can usually be done in 1-2hrs from opening the ticket. For server issues; again, pretty easy. Just use iKVM/IPMI and iPXE to diagnose a faulty server. Again, using "remote-hands" from the CoLo provider can help fix problems if your staff does not have the skills…

In my experience, the issues that take 80% of your time are the unexpected edge cases, not the easy fixes.

Swapping drives is basically the easiest fix. The issues that cause the most problems are the hard to diagnose ones like the faulty RAM that flips a bit every once in a while or the hard drive controller that triggers an driver bug with weird behavior that doesn’t show up in the logs with anything meaningful.

Re: Building the heap: racking 30 petabytes of hard drives for pretraining

#258
post #199
post #187

Earlier quoted context omitted.

Massive props for getting it done anyway. For others reading: In general a switch should never run DHCPd, but will normally/often relay it for you, your arista's would 100% have supported relaying, but in this case it sounds like it might even be flat L2. Normally you'd host dhcpd on a server. Some general feedback incase it's helpful.. -20K on contractors seems insane if we're talking about rack and stack for 10 rac…

def agree on the setup fees, that was just a price crunch to get it done within the weekend. (too short-notice for professional services, too sensitive for craigslist, so basically just paying a bunch of folks we already knew and trusted) for IPXE do you have any reference material you'd recommend? we had 3 people each with reasonably substantial server experience try for like 6 hours each and for whatever reason it…

Similarly I'm happy to share my ipxe scripts. It's just one of those things that you need to understand the fundamentals of before you start. It's about a hundred lines of bash to setup.

Re: Building the heap: racking 30 petabytes of hard drives for pretraining

#259
post #232
post #186

Earlier quoted context omitted.

Rackspace is typically at a premium at most data centers.

My info may be dated, but power density has gone up a ton over time. I'd expect a lot of datacenters to have plenty of space, but not much power. You can only retrofit so much additional power distribution and cooling into a building designed for much less power density.

This is my experience as well. We have 42u racks with 8 machines in them because we cant get more power circuits to the rack.

Re: Building the heap: racking 30 petabytes of hard drives for pretraining

#260

Well done! I love the honest write up and the “can do” attitude. Must have been a lot of fun too. Out of interest why do you think you made the mistake of buying 20x more drives than you needed instead of the denser storage that you mention? Was there a reason you opted for this?

He did mention that it would have been a higher up-front cost.
Post reply on HN