I imagine Mr Good Guy at OVH telling some others: "guys we have a single point of failure in our architecture with SBG, maybe we should... - naaah it's fine, we do not have time nor resources" Then shit happens. edit: I have no idea what is happening exactly, but OVH being what it is, it seems extremely weird that all datacenters "can" get down at the same time, and it looks like a serious architecture problem to me…
OVH Incident in Strasbourg
151–160 of 207 posts
Re: OVH Incident in Strasbourg
#152I imagine Mr Good Guy at OVH telling some others: "guys we have a single point of failure in our architecture with SBG, maybe we should... - naaah it's fine, we do not have time nor resources" Then shit happens. edit: I have no idea what is happening exactly, but OVH being what it is, it seems extremely weird that all datacenters "can" get down at the same time, and it looks like a serious architecture problem to me…
> - naaah it's fine, we do not have time nor resources" Yup, been there multiple times in smaller hosting companies. It's basically how it goes. They don't get serious about outages until revenue is severely affected and the brand damaged, they don't get serious about security until there's been a big breach or sales are lost because of lack of certification.
Re: OVH Incident in Strasbourg
#153Maybe their main DCs, or their largest, but not all of them. I have virtual servers in thier Quebec DC (BHS) and it hasn't gone down since the last time I rebooted it.
Re: OVH Incident in Strasbourg
#154I imagine Mr Good Guy at OVH telling some others: "guys we have a single point of failure in our architecture with SBG, maybe we should... - naaah it's fine, we do not have time nor resources" Then shit happens. edit: I have no idea what is happening exactly, but OVH being what it is, it seems extremely weird that all datacenters "can" get down at the same time, and it looks like a serious architecture problem to me…
It was a power outage. Their own generators did not work. This explain why all is down, but there is surely something bad in their architecture.
EDIT: The title on HN in misleading, summary from their CEO here – https://twitter.com/olesovhcom/status/928592231807713280
Re: OVH Incident in Strasbourg
#155Earlier quoted context omitted.
While at $bigco we halted testing of generation equipment because it was sending DCs offline more often than it kept them up. Lawyers were involved, things got ugly
I'm completely unfamiliar with electrical generators/power generation, so take this question in the spirit of ignorance: Is there not a way to test generators without actually having them power the live datacenter infrastructure? I mean, simulate the exact generation and load requirements that the generators will face? I don't know if it's feasible to dump all that power to ground or whatever, but that way you could…
You know all that cooling equipment data centers tend to have?
Dumping power to ground is also known as an electric furnace.
Now you have twice as much heat to move, but you only have the usual cooling system. Toasty servers are sad servers. Toasty engineers are dead engineers.
Re: OVH Incident in Strasbourg
#156Re: OVH Incident in Strasbourg
#157Earlier quoted context omitted.
People not actually testing emergency equipment.
While at $bigco we halted testing of generation equipment because it was sending DCs offline more often than it kept them up. Lawyers were involved, things got ugly
Re: OVH Incident in Strasbourg
#158Earlier quoted context omitted.
What is the point of backup generators if you do not verify that they work every so often? I have a very hard time believing that they actually tested that they worked, because a failure of not one, but both of them.
To be fair people do test generators on a monthly schedule usually. Problem you find is it’s getting colder now so any problems are amplified suddenly. Might have been entirely tested a couple of weeks ago.
There was a DC in California a few years back that had three generators fail during a scheduled test and one of their server rooms had a blackout because of it.
Re: OVH Incident in Strasbourg
#159To make error is human. To propagate error to all server in automatic way is #devops. - @devopsborat
just in case you followed the wrong devops borat, it's @DEVOPS_BORAT (the other one is a spam bot).
@DNS_BORAT https://twitter.com/DNS_BORAT
@InfoSecBorat https://twitter.com/InfoSecBorat
@KanbanBorat https://twitter.com/KanbanBorat
@mysqlborat https://twitter.com/mysqlborat
@NetEng_Borat https://twitter.com/NetEng_Borat
@secure_borat https://twitter.com/secure_borat
@SecurityBorat https://twitter.com/SecurityBorat
@Sysadm_Borat https://twitter.com/Sysadm_Borat
Re: OVH Incident in Strasbourg
#160Damn, every emergency power supply I have encountered (the big ones with fuel and hundreds of batteries) always fail to start when they have to... Why is that ?
People not actually testing emergency equipment.
For example you switch off the data centre circuit breakers and everything fails over to generators just fine. Test successful, right?
Then when there's a real outage you have problems because the operations team's computers have all gone off, so they can't migrate load to a different data centre. It didn't happen in testing, because they aren't in the data centre so their kit isn't connected to the breakers you turned off.
Or it turns out the wireless APs aren't on UPSes. Or it turns out there's a switch in a closet somewhere that isn't on a UPS. Or they tested for a single loss of power, but when the mains power toggles on and off every 30 seconds the UPS batteries get run down. Or they need to top up the generator and they discover you can't get fuel delivered at 9pm on a Friday. Or the generator doesn't recharge the UPS, but you have to turn off the generators to refuel them. Or a guy had a standalone UPS for his desktop, but his monitor wasn't connected as the UPS only came with IEC C13 power cables and his monitor needed IEC C5...