They now make their own if I remember correctly.
https://image.slidesharecdn.com/cpn208failuresatscale-121129...
21–30 of 129 posts
They now make their own if I remember correctly.
https://image.slidesharecdn.com/cpn208failuresatscale-121129...
I have a hard time understanding why they needed humans babysitting the restart of servers and services. Servers and services are supposed to automatically restart after the power came back. This implies OVH doesn't regularly test server restarts with something like Netflix Chaos Monkey. Moreover, the power lines were not really redundant in SBG, and regarding the network downtime in RBX, it looks like the network co…
OVH runs servers for their clients, do you really expect them to randomly restart their clients' servers? Netflix is not a hosting provider, the situations are not comparable.
Why are backup generators so unreliable? And that is unreliable by the standard of a non-critical system let alone that of an emergency system where reliability is their sole purpose. Is it just a (non)-survivorship bias where we most commonly talk about the failure cases and they actually have a stellar 99.99999% record that doesn't make headlines?
I think they're just large, complicated physical things where stuff goes wrong, even if you're maintaining and testing. My best story was as a young student turning up for my helpdesk shift at about 5:50am to a phalanx of fire engines and the Hazmat team. The generator in the basement was maintained and tested. This time it had started when power went out, but there was a pump that filled a holding tank from a 10,000…
I have a hard time understanding why they needed humans babysitting the restart of servers and services. Servers and services are supposed to automatically restart after the power came back. This implies OVH doesn't regularly test server restarts with something like Netflix Chaos Monkey. Moreover, the power lines were not really redundant in SBG, and regarding the network downtime in RBX, it looks like the network co…
OVH runs servers for their clients, do you really expect them to randomly restart their clients' servers? Netflix is not a hosting provider, the situations are not comparable.
Very transparent, kudos for that. A reminder to everyone, though, that the brochure doesn't always match reality. OVH pitches itself as a global data center leader, with 20 facilities, 2 of them amongst the largest in the world. Not to harp on them, just to say that having some backup plan on a separate provider is always a good idea, no matter who your primary is.
Also when you decide on a backup, make sure your backup is not hosting their servers in the same DCs as is your primary, or worse yet that your backup is renting their servers from the same people you rent your primary from.
Earlier quoted context omitted.
Most of the time, it's not the actual generator that fails. It's often either the part that detects power failure (fluctuations, brownouts, ... are harder to detect than total failure) or the switchover hardware to the generator. If you want to switch 20KV, you're really handling an awful lot of energy. Physics kinda get in the way there. Switching paket based networks is often much easier because you can just hold o…
To add to this, backup power systems are in fairly widespread deployment and bigger installations are tested and maintained on a schedule. You'll never see "hospital transferred emergency-available circuits to backup system; no malfunction" in the news, yet it happens every few weeks in every installation. Generators themselves are very reliable machines. The engines are cast iron (instead of cast aluminium as found…
Earlier quoted context omitted.
OVH runs servers for their clients, do you really expect them to randomly restart their clients' servers? Netflix is not a hosting provider, the situations are not comparable.
I would be pissed if my hosting provider randomly restarted my servers, taking them offline for minutes each time.
Earlier quoted context omitted.
To add to this, backup power systems are in fairly widespread deployment and bigger installations are tested and maintained on a schedule. You'll never see "hospital transferred emergency-available circuits to backup system; no malfunction" in the news, yet it happens every few weeks in every installation. Generators themselves are very reliable machines. The engines are cast iron (instead of cast aluminium as found…
Apparently OVH did test their generator system but the outage detector or controller failed.
If generators really are just a giant blob of metal and very simple (and I have no reason to disbelieve the GP comment)... well... it could be kind of interesting to build an open source software stack to handle switchover. Because, disclaimers notwithstanding, the code would ostensibly be super simple too. So even if it couldn't officially be used directly, it would certainly provide a good base for engineers to copy-paste, either literally or ideologically (and then thoroughly verify, of course!).
Okay... thinking about it, I'm probably wrong - either generator control has some fundamental intricacies that make it not-completely-simple, or all sites have edge cases that have to be baked into the firmware by a system integrator/electrical engineer.
I say this because I'm (genuinely) trying to figure out why the PLCs failed - both in OVH's case, and in AWS's case (see elsewhere in this thread, https://news.ycombinator.com/item?id=15676189).
It's obvious there are crazy but legit reasons for this kind of thing to happen. I'm very curious what the complexity scale here is.
Very transparent, kudos for that. A reminder to everyone, though, that the brochure doesn't always match reality. OVH pitches itself as a global data center leader, with 20 facilities, 2 of them amongst the largest in the world. Not to harp on them, just to say that having some backup plan on a separate provider is always a good idea, no matter who your primary is.
Also when you decide on a backup, make sure your backup is not hosting their servers in the same DCs as is your primary, or worse yet that your backup is renting their servers from the same people you rent your primary from.
The reason was problem with one of Sprint's cables. It was carrying a lot of Nordunet traffic. Nordunet is/was the Nordic joint university network infrastructure. At the time, for historical reasons, Nordunet had more network capacity than most of the rest of Europe combined (consider that Norway was the second international connection to Arpanet in the early 70's due to the seismic array at Kjeller outside Oslo that provided essential data on Soviet nuclear tests; this made money flow to network research in Norway, and a joint Nordic effort let to significant investment in the other Nordic countries too; couple it with early work in Sweden, where they ended up hosting D-Gix - for many years the largest internet exchange point outside the US).
So Sprints cable went down, and everyone started re-routing. Problem #1: Everyone had high capacity to D-Gix or other exchange point in the Nordic countries. These connections were now flooded by traffic from the Nordic networks where people were used to being able to saturate their 100Mbps network connections (it was such a downer starting my ISP, and going from 100Mbps at university to 512kbps aggregate outbound capacity at our offices). Except most of the links connecting to D-Gix where 34Mbps or less, and there were not that many of them, and they were connecting entire universities or entire countries...
So connectivity within Europe started slowing down.
Problem #2: Most of the links from elsewhere in Europe to the US were 8Mbps or less, with maybe a couple ones faster than that, and there were not that many of them. Certainly not enough to compensate for several hundred Mbps of lost capacity.
Problem #3: Many of those links got so saturated and slow that failsafes were tripped and traffic started getting routed to their expensive backup links (e.g. people paying for a port and then paying for burst a 95th percentile basis and the like with expectation of normally not using them). Except, as it turns out, many of said backup links had been purchased from Sprint. Over that cable.
Today the number of cables and overall capacity is vastly higher, with many more providers, so I don't think the same could happen again, and D-Gix is no longer as important as it was (though it's still one of the largest international interchange points), but as we sat there pinging and trying to figure out what was up, it was a vivid lesson in ensuring your backups are actually sufficiently separate from your main systems that they stand a chance of working.
Even if we lost power on one power channel, we would still be operational. Even if the backup power failed. That's the reason you use redundant power.
If your "cloud provider" is charging you a lot but not providing your server redundant power from two grids to redundant PSUs, I don't think you're getting the best available design.
In this case, it looks like OVH just threw all the power onto a single channel, which allowed a failover bug to bring down the datacenter.