Live data from Hacker News

OVH Incident in Strasbourg

status.ovh.com

191–200 of 207 posts

Re: OVH Incident in Strasbourg

#191
post #105

Earlier quoted context omitted.

People not actually testing emergency equipment.

I've seen places that /did/ test their backup power - but they got failures anyway, because of faults the test didn't reveal. For example you switch off the data centre circuit breakers and everything fails over to generators just fine. Test successful, right? Then when there's a real outage you have problems because the operations team's computers have all gone off, so they can't migrate load to a different data cen…

One place I worked at found this during the big storm in the UK UPS worked fine and all Telecoms Golds Machines stayed - but they had forgotten to put the Modems on the UPS :-)

Re: OVH Incident in Strasbourg

#192
post #136
post #105

Earlier quoted context omitted.

People not actually testing emergency equipment.

While at $bigco we halted testing of generation equipment because it was sending DCs offline more often than it kept them up. Lawyers were involved, things got ugly

Sounds like your doing it on the cheap

Re: OVH Incident in Strasbourg

#193
post #136

Earlier quoted context omitted.

While at $bigco we halted testing of generation equipment because it was sending DCs offline more often than it kept them up. Lawyers were involved, things got ugly

I'm completely unfamiliar with electrical generators/power generation, so take this question in the spirit of ignorance: Is there not a way to test generators without actually having them power the live datacenter infrastructure? I mean, simulate the exact generation and load requirements that the generators will face? I don't know if it's feasible to dump all that power to ground or whatever, but that way you could…

Your test would only test the generators. It would not test the transfer switches, or whether everything you think is connected to the backup generator is actually connected, and any equipment in between.

Re: OVH Incident in Strasbourg

#194
post #182

Earlier quoted context omitted.

This happens all the time. Every single thing that you see in software development happens in network engineering and data center engineering, except that where in development in general senior people who write software are capable of at least guestimating complexities to provide a marginally unified front against the unreasonable expectations of execs, it is pretty much never the case in neteng or dcops as those tha…

Any decent sr. network engineer or architect should be able to design you a network and explain the pros and cons, risks, and future scalability. If someone doesn't know why a failure occured and they can't find out then they aren't looking hard enough

Please tell me about about this fabulous senior network engineer and where I can obtain a dozen of them.

Re: OVH Incident in Strasbourg

#195

The status page is up again http://status.ovh.net/ I paste the report so far: ------------- FS#15162 — SBG Attached to Project— Network Task Type: Incident Category: Strasbourg Status: In progress Percent Complete: 0% Details We are experiencing an electrical outage on Strasbourg site. We are investigating. Comments (2) Comment by OVH - Thursday, 09 November 2017, 10:55AM SBG: ERDF repared 1 line 20KV. the second is…

More:

Comment by OVH - Thursday, 09 November 2017, 12:44PM

Everything is back up electrically. We are checking that everything is OK and we are identifying still impacted services/customers.

Comment by OVH - Thursday, 09 November 2017, 13:25PM

Hello, Two pieces of information,

This morning we had 2 separate incidents that have nothing to do with each other. The first incident impacted our Strasbourg site (SBG) and the 2nd Roubaix (RBX). In SBG we have 3 datacentres in operation and 1 under construction. In RBX, we have 7 datacentres in operation.

SBG: In SBG we had an electrical problem. Power has been restored and services are being restarted. Some customers are UP and others not yet. If your service is not UP yet, the recovery time is between 5 minutes and 3-4 hours. Our monitoring system allows us to know which customers are still impacted and we are working to fix it.

RBX: We had a problem on the optical network that allows RBX to be connected with the interconnection points we have in Paris, Frankfurt, Amsterdam, London, Brussels. The origin of the problem is a software bug on the optical equipment, which caused the configuration to be lost and the connection to be cut from our site in RBX. We handed over the backup of the software configuration as soon as we diagnosed the source of the problem and the DC can be reached again. The incident on RBX is fixed. With the manufacturer, we are looking for the origin of the software bug and also looking to avoid this kind of critical incident.

We are in the process of retrieving the details to provide you with information on the SBG recovery time for all services/customers. Also, we will give all the technical details on the origin of these 2 incidents.

We are sincerely sorry. We have just experienced 2 simultaneous and independent events that impacted all RBX customers between 8:15 am abd 10:37 am and all SBG customers between 7:15 am and 11:15 am. We are still working on customers who are not UP yet in SBG. Best, Octave

Re: OVH Incident in Strasbourg

#196
post #16

I moved away from OVH after I paid 3 months advance (~$300) for a server which burned down after 1 1/2 months. They did not issue any refunds (data, blood, sweat and tears were lost that day). I have been an OVH customer for 12 years. Today, I'm glad to have moved away all my production environments as well.

Whatever.

Servers break. Providers go down.

It happens to all of them. Whenever I read comments like yours I wonder where you moved your servers to, and if you'll move them again when that goes down too.

Re: OVH Incident in Strasbourg

#197

Earlier quoted context omitted.

Ah, gotcha. So, that one datacenter caused all other datacenters to die..?

It's supposed to be two separate incidents: power going down in Strasbourg, and fiber network equipment going down in Roubaix (the main center of OVH's network) due to a "software bug". It's explained here https://twitter.com/olesovhcom/status/928587258583748609 in French, they might post an English-language translation soon.

Thanks!

Re: OVH Incident in Strasbourg

#198
post #182

Earlier quoted context omitted.

This happens all the time. Every single thing that you see in software development happens in network engineering and data center engineering, except that where in development in general senior people who write software are capable of at least guestimating complexities to provide a marginally unified front against the unreasonable expectations of execs, it is pretty much never the case in neteng or dcops as those tha…

Any decent sr. network engineer or architect should be able to design you a network and explain the pros and cons, risks, and future scalability. If someone doesn't know why a failure occured and they can't find out then they aren't looking hard enough

I think I'm a pretty decent senior network engineer but I've been hit by firmware bugs more than once. No matter how much redundancy you build in, there's always shit that can go wrong and there will always be things that you just can't foresee (like the bug in this case). The software on these network devices and optical gear is written by humans, like all software, and is not perfect, like all software.

(In one case I experienced, the vendor reassured me that what I reported was not even technically possible -- until one of their engineers flew out and witnessed it firsthand.)

Re: OVH Incident in Strasbourg

#199
post #109

Earlier quoted context omitted.

>ts OVH: The Hardware is good DDoS protection is good The Prices are high The prices are high? Compared to what? Cheap is their raison d'être.

compared to other dedicated Server Hardware, not talking about business Cloud Infrastructure, no idea about that. Sry if that caused confusion

I was considering their entire line...soyoustart, kimsufi, and ovh. There's not much cheaper than them. Hetzer in some cases, but not all.

Re: OVH Incident in Strasbourg

#200

Earlier quoted context omitted.

I'm completely unfamiliar with electrical generators/power generation, so take this question in the spirit of ignorance: Is there not a way to test generators without actually having them power the live datacenter infrastructure? I mean, simulate the exact generation and load requirements that the generators will face? I don't know if it's feasible to dump all that power to ground or whatever, but that way you could…

You can if you have to. But then you're really only doing a fancy simulation. I accompanied my dad (power engineer) to a water purification plant where they were testing new equipment for the back up generator. There their weekly tests involved moving the entire plant to the diesel generator and running it of back up power for a couple of hours (once you start a big generator you have to let it run or it wont last lo…

P.S. Every test is a simulation of reality. At Fukushima the diesel generators flooded. Lesson - the unknown reason that'll knock out your grid can knock out your backup

well the lesson there was more like, that it's stupid to put your diesel generators deep in the ground when they should sustain burst sea level raises. (well I think it's never a good idea to do that, I've seen special places to put them even deep inside germany, just because some panicful people that might think that it still could overflow with ground water, etc)

Post reply on HN