Live data from Hacker News

OVH CEO Octave Klaba speaking about the incident [video]

ovh.com

131–140 of 143 posts

Re: OVH CEO Octave Klaba speaking about the incident [video]

#131
post #65

Earlier quoted context omitted.

This is a difference in what you are buying. When you are buying a dedicated server, there isn't exactly a good way to hide that the thing has just gone up in smokes. When you buy a storage API, sure, failure rates go up, latency increases 100x, but after a few hours its probably back to normal. Of course, with the increased abstraction, you get more problems. "Availability zones" are useless when most cloud outages…

AWS is the new IBM. Nobody ever got fired for using AWS.

And funny enough, it's now considered crazy to do a hardware startup, even if staffed up with industry vets. The reaction to Oxide is funny to watch, especially.

Re: OVH CEO Octave Klaba speaking about the incident [video]

#132

Earlier quoted context omitted.

> The problem with water in the same place with high power equipment […] Most high-end UPSes have a relay where you can run an active-high or active-low emergency power off (EPO) signal. The EPO can either be a button that is pressed manually by the staff, automatically via fire suppression system, or both/either. Schneider-APC white paper (PDF) * https://download.schneider-electric.com/files?p_File_Name=AS... The EP…

Such a switch doesn't make any remaining battery charge in the UPS go away magically. If the UPS housing gets breached (e.g. ingress of impure water or bending/melting from heat), you're back to square one.

Batteries can be held in a different room than the IT equipment and UPS inverters--up to 200m per the manual of the APC Galaxy VX.

So can have the more common water fire suppression in the day-to-day areas where people are more likely to be in, and have have non-water solutions. Water mist instead of 'traditional' sprinklers is also a thing in many places:

* https://res.mdpi.com/d_attachment/energies/energies-13-05117...

Re: OVH CEO Octave Klaba speaking about the incident [video]

#133
post #67

Earlier quoted context omitted.

Not halon any more, but instead an ozone-friendly fire suppression gas such as argonite

That's basically true but I want to note that halon is a gas that actually suppresses fire while argonite is just there to get rid of oxygen.

Yes, one of the big differences between the old halon system we got rid of many years ago, and the newer argonite systems, is that the argonite requires many more gas cylinders, because (as you say) argonite works by getting rid of oxygen, instead of chemically suppressing combustion. The small machine room in our office building was refitted about 8 years ago with argonite fire suppression, and they had to add pressure release vents between the machine room and the outside so that there was somewhere for a room full of air to be displaced to.

AIUI being in a room during a halon dump is unpleasant but not too dangerous, but you really don’t want to be in there when the argonite release happens!

Re: OVH CEO Octave Klaba speaking about the incident [video]

#134

Earlier quoted context omitted.

And Facebook had small UPS/ATS units at the end of each row. Not sure if they still do that today it was like that when I walked through their datacenter. They did that for the purpose of power efficiency. They lost far less power by having many smaller units.

What you saw likely wasn't a UPS, but an RPP > A data center typically spans four rooms, called suites, where racks of servers are arranged in rows. Up to four MSBs provide power to each suite. In turn, each MSB supplies up to four 1.25 MW Switch Boards (SBs). From each SB, power is fed to the 190 KW Reactive Power Panels (RPPs) stationed at the end of each row of racks. https://research.fb.com/wp-content/uploads/201…

They specifically called em out as UPS and talked about the efficiency improvements over having a dedicated UPS room. I don't remember the percentage numbers though. It was orders of magnitude less power loss. It could be they ditched all of that hardware by now. This was some time ago. You might ask some of the old timers in the DC. They probably have pictures of the hardware.

Re: OVH CEO Octave Klaba speaking about the incident [video]

#136
post #65

This would be one of inherient difference between smaller vs. giga players in cloud hosting. AWS/Google/Azure, if this happens, there should only be limited outage to a small fraction of customers. As a matter of fact, Google had such an incident before, and literally no customers (internal and external) noticed.

This is a difference in what you are buying. When you are buying a dedicated server, there isn't exactly a good way to hide that the thing has just gone up in smokes. When you buy a storage API, sure, failure rates go up, latency increases 100x, but after a few hours its probably back to normal. Of course, with the increased abstraction, you get more problems. "Availability zones" are useless when most cloud outages…

If you're in with the big cloud providers you have no choice. Hybrid cloud is economically impossible due to the bandwidth costs.

Yet somehow, at smaller providers and dedicated hosters bandwidth is usually included as a too-cheap-to-meter feature. Gotta love cloud innovation.

Re: OVH CEO Octave Klaba speaking about the incident [video]

#137
post #95
post #84

Earlier quoted context omitted.

> oxygen replacing suppressants No, they are a major risk to the employees working there. I would much rather have a data center destroyed by for every twenty years without victims than mandating the user of oxygen replacing fire suppressants.

My old company's DC uses FM200 gas which puts fires out without asphyxiating anyone (I'm assuming the experience would be a respiratory workout for anyone caught in it, but it would have cost £20000 to test).

Halon 1301 is still used for fire suppression aboard ships and aircraft. It can be safely flooded into confined spaces without killing people.

Calling these agents "oxygen suppressing" is a HUGE misnomer. Halon, FM200, etc don't work by reacting with the oxygen in the room. Although displacing some oxygen contributes to their method of action, this isn't the primary way they put out fires.

Halon and friends stop fires by catalytically interrupting gaseous fire reaction products. As it was explained to me, these reactions can be counterintuitive - eg, the production of reduced hydrogen gas (H2) from free radicals. You wouldn't expect Halon to put out a fire by making hydrogen, but it does.

Re: OVH CEO Octave Klaba speaking about the incident [video]

#138
post #15

Is it true that OVH has literally all data in that single datacenter?

Genuine question: How did that question come to be? Knowing nothing about OVH, I just typed "ovh datacenters" into Google and the first hit was this: https://www.ovh.com/world/us/about-us/datacenters.xml with the first sentence being "27 data centers around the world, including 2 of the largest ones".

weird '.xml' url.. wonder where that's linked from

list broken down a bit more on this page

https://us.ovhcloud.com/about/company/data-centers

Re: OVH CEO Octave Klaba speaking about the incident [video]

#139

Earlier quoted context omitted.

Just to note - only google does live migration (with ~100ms blackout) with others you will take some downtime anyway (assuming the remaining zones even have enough capacity to fit everyone)

Google never migrates to a different physical location, so it wouldn't actually protect against a whole class of issues (major flooding, war, employee strike, etc)

Zones are often in separate locations but only within a few miles so yeah. Also i think they won’t even migrate inter-zone so there’s that

Re: OVH CEO Octave Klaba speaking about the incident [video]

#140
post #35

Earlier quoted context omitted.

Oh good to know. I don't use OvH, and my limited understanding were from their products page which lists VM style offerings. I had assumed VMs were the major use cases on OvH.

The pricing on their dedicated servers are cheaper than most 1-2GB VMs on cloud: https://www.ovhcloud.com/en/bare-metal/ - this is their flagship and most expensive brand, the cheaper ones are even less The funniest tweets demanding their data and saying they'll lose everything are the people running: https://twitter.com/Sensity_RP/status/1369496048998223873 - GTA5 multiplayer gameserver that begs for donations. Runn…

I can imagine a provider that offers "Disaster Recovery Plan" in the sense that you push a button and they spin up [whatever machines you subscribe to] in a different data center and restore your latest backup to it.

Or the button flips your DNS from your usual primary servers to a hot backup (or a cold backup that you brought online yourself).

Post reply on HN