Live data from Hacker News

OVH CEO Octave Klaba speaking about the incident [video]

ovh.com

81–90 of 143 posts

Re: OVH CEO Octave Klaba speaking about the incident [video]

#81

Are there industry options or methods of wiring to allow for a UPS room separate from the actual rooms the racks are stored in? It's almost tradition to have a rack with UPS's in the bottom and then the rest of the space filled with servers or drive arrays. We wouldn't ever think of putting a tiny backup generator in the bottom of every rack, so why do we put a battery storage system there? Also, with the advances in…

For the extreme opposite of that, Google famously trolled everyone 10 years ago by announcing that every one of their servers had its own in-chassis 12V battery: https://www.cnet.com/news/google-uncloaks-once-secret-server...

It’s true! I wrote software that upgraded firmware on every one of those batteries without frying them up (most of the time). There was a public paper/talk on that few years back

Re: OVH CEO Octave Klaba speaking about the incident [video]

#82
post #33

This would be one of inherient difference between smaller vs. giga players in cloud hosting. AWS/Google/Azure, if this happens, there should only be limited outage to a small fraction of customers. As a matter of fact, Google had such an incident before, and literally no customers (internal and external) noticed.

It also depends if you're renting a dedicated server, vs cloud/VPS. AWS/Google/Azure deal with virtualized systems that can be moved around to another server easily. OVH has a lot of dedicated servers as well though, so if you're using one of those then it can't be moved very easily to avoid downtime.

Just to note - only google does live migration (with ~100ms blackout) with others you will take some downtime anyway (assuming the remaining zones even have enough capacity to fit everyone)

Re: OVH CEO Octave Klaba speaking about the incident [video]

#83
post #59
post #58

Earlier quoted context omitted.

2 OVH dedicated servers in different countries are still cheaper than one AWS instance. e.g. I've got servers at OVH SBG-2, Hetzner's Falckenstein, and Online.net's AMS datacenters — the total of which is still almost a magnitude less than the same cost on AWS or GCP (granted, that's including traffic)

Pretty much. Use 5% of the money you saved moving your workload from AWS to OVH to support failover, DR, and backups. You can probably buy like five or six machines of equivalent spec or more.

Not only that, but the OVH load balancer (API gateway) is enormously much cheaper at 2TB of traffic than AWS and they don't charge for the amount of requests.

Re: OVH CEO Octave Klaba speaking about the incident [video]

#84
post #24

I think the key issue here is that there wasn't a functioning fire suppresssion system in place. Second question: is such a system required for this kind of operation? Maybe?

If you're running an IT operation this big, you should have both fire detection units and oxygen replacing suppressants. That should be mandatory. Otherwise, it'd be very hard to contain. Especially when all your servers have RAID controllers with Li-Ion batteries or supercapacitors or other extremely trigger-happy components. Oh, and cooling systems. You're just kindling the fire with it at the beginning.

> oxygen replacing suppressants

No, they are a major risk to the employees working there.

I would much rather have a data center destroyed by for every twenty years without victims than mandating the user of oxygen replacing fire suppressants.

Re: OVH CEO Octave Klaba speaking about the incident [video]

#85
post #84

Earlier quoted context omitted.

If you're running an IT operation this big, you should have both fire detection units and oxygen replacing suppressants. That should be mandatory. Otherwise, it'd be very hard to contain. Especially when all your servers have RAID controllers with Li-Ion batteries or supercapacitors or other extremely trigger-happy components. Oh, and cooling systems. You're just kindling the fire with it at the beginning.

> oxygen replacing suppressants No, they are a major risk to the employees working there. I would much rather have a data center destroyed by for every twenty years without victims than mandating the user of oxygen replacing fire suppressants.

I understand the concerns you have, however I think there's a good middle ground. Would you consider the following procedure acceptable?

    - Isolate all rooms with fire-proof doors.
    - Keep fire supression system at manual.
    - When fire breaks try to contain (we have 24h watch).
    - If fails trigger fire supression system. It has 90 second delay and activated per room.
    - Leave premsises, make the calls.
Fire control and supressant control is not inside the system room. Also fire-proof doors seal the room reasonably well, so chemical doesn't move freely. Also there are better chemicals which break down faster and less harmful to everything.

BTW, We use Novec.

Re: OVH CEO Octave Klaba speaking about the incident [video]

#86

Earlier quoted context omitted.

I’m pretty confident they have efficient fire suppression systems. They are hosting at least 400,000 servers. They for sure have multiple servers taking fire every single day, and yet it’s the first time it ends up in a catastrophic fire. The fire suppression system either catastrophically failed, or something out of design happened with one of their inverters, as suggested in the video.

Do servers really catch on fire nearly daily at this scale?

Yes. But normally it doesn't get outside the chassis of the machine.

Normally upon disassembly you'll find a few burnt components and a sooty mark the size of your fist.

Re: OVH CEO Octave Klaba speaking about the incident [video]

#87
post #33

Earlier quoted context omitted.

It also depends if you're renting a dedicated server, vs cloud/VPS. AWS/Google/Azure deal with virtualized systems that can be moved around to another server easily. OVH has a lot of dedicated servers as well though, so if you're using one of those then it can't be moved very easily to avoid downtime.

Just to note - only google does live migration (with ~100ms blackout) with others you will take some downtime anyway (assuming the remaining zones even have enough capacity to fit everyone)

Google never migrates to a different physical location, so it wouldn't actually protect against a whole class of issues (major flooding, war, employee strike, etc)

Re: OVH CEO Octave Klaba speaking about the incident [video]

#88
post #22

This would be one of inherient difference between smaller vs. giga players in cloud hosting. AWS/Google/Azure, if this happens, there should only be limited outage to a small fraction of customers. As a matter of fact, Google had such an incident before, and literally no customers (internal and external) noticed.

I can't even find any press articles about the Google incident.

At the time that Google data center was not known to be owned by Google. So it will just be "fire in industrial estate"

Re: OVH CEO Octave Klaba speaking about the incident [video]

#89
post #84

Earlier quoted context omitted.

> oxygen replacing suppressants No, they are a major risk to the employees working there. I would much rather have a data center destroyed by for every twenty years without victims than mandating the user of oxygen replacing fire suppressants.

I understand the concerns you have, however I think there's a good middle ground. Would you consider the following procedure acceptable? - Isolate all rooms with fire-proof doors. - Keep fire supression system at manual. - When fire breaks try to contain (we have 24h watch). - If fails trigger fire supression system. It has 90 second delay and activated per room. - Leave premsises, make the calls. Fire control and su…

Other industries use lock-out keys when people are in harm's way. It seems like it'd be easy enough to design an oxygen-replacement system that has lock-outs that people engage whenever they need to enter the protected rooms.

https://en.wikipedia.org/wiki/Lockout–tagout

Re: OVH CEO Octave Klaba speaking about the incident [video]

#90
post #29

Are there industry options or methods of wiring to allow for a UPS room separate from the actual rooms the racks are stored in? It's almost tradition to have a rack with UPS's in the bottom and then the rest of the space filled with servers or drive arrays. We wouldn't ever think of putting a tiny backup generator in the bottom of every rack, so why do we put a battery storage system there? Also, with the advances in…

There are plenty of datacenters with a separate battery room, sure.

So far I have never been in one where this was not the case. Also they always have fire suppression.
Post reply on HN