Live data from Hacker News

OVH Incident in Strasbourg

status.ovh.com

91–100 of 207 posts

Re: OVH Incident in Strasbourg

#91
post #22

It started with all our SBG servers going down simultaneously. Approximately 1h later all our RBX servers went down as well including the OVH status page and all other OVH web applications. Either their SBG and RBX data centers are somehow connected or those are indeed two independent incidents.

SBG was a power failure, followed by a generator failure. Not sure if they set up their network infrastructure in a way that this could spread but I find it hard to imagine that the outage in SBG triggered RBX going completely black. Unless of course they store their configuration files in SBG.

It seems RBX was an unrelated failure corrupting the router configuration (@olesovhcom).

SBG going down due to a quadruple power failure (both grid connections and both generators) is quite spectacular.

Re: OVH Incident in Strasbourg

#92

Title is misleading. Only RBX and SBG were affected. 06:15 UTC SBG serves failed. OVH network weathermap: http://weathermap.ovh.net Btw. First post: https://news.ycombinator.com/item?id=15660524

My system in Lille has also been up without problems, at least since I got in at work.

Re: OVH Incident in Strasbourg

#93

Network and RBX are UP again: https://twitter.com/olesovhcom/status/928556358353539072 (but SGB's datacenters are still being restarted)

No it isn't.

Servers appear up just the website is struggling (probably everyone logging in at the same time to file a ticket and complain).

Re: OVH Incident in Strasbourg

#94
post #46
post #39

Earlier quoted context omitted.

In this case, I wouldn't be to hard on them. As it appears they lost their main power line, the backup power line and both generators failed and one generator has been restarted now.

So what? Losing main power is a standard case for any DC. That's why you have generators. Even a generator failure is nothing out of the ordinary. But that no generators in a DC work kind of indicates that they don't test them as often as you would expect. They just announced that they want to be a "hypercloud" provider on the scale of AWS and Google Cloud. I really hope that a power failure in Virginia couldn't brin…

It doesn't matter how much testing you do – unfortunately some things can still fail.

I'm much more interested in why the failure cascaded to other data centres – that's exactly what shouldn't happen.

Re: OVH Incident in Strasbourg

#95
post #53

Earlier quoted context omitted.

To be fair people do test generators on a monthly schedule usually. Problem you find is it’s getting colder now so any problems are amplified suddenly. Might have been entirely tested a couple of weeks ago.

If your generators are not reliable in the cold, the issue is the placement of them. This is something basic to account for.

The problem is usually not "not reliable in cold" but rather, the generators are X years old and the temperature is now changing in Europe from "mostly warm" to "warm over the day and icecold in the night" and finally aiming for "icecold all day", which means any equipment exposed will go through rather severe temperature changes.

While generators are usually able to handle this with sufficiently low failure risk, the risk is increase due to the changing temperature

Re: OVH Incident in Strasbourg

#96
post #53

Earlier quoted context omitted.

To be fair people do test generators on a monthly schedule usually. Problem you find is it’s getting colder now so any problems are amplified suddenly. Might have been entirely tested a couple of weeks ago.

If your generators are not reliable in the cold, the issue is the placement of them. This is something basic to account for.

Right, but that was obviously just an example.

Their generators are almost certainly tested frequently. But there could be any number of causes underlying the failure, and unfortunately sometimes failure does happen.

Re: OVH Incident in Strasbourg

#97
post #83

I had never heard of this company til I saw this post. Shrugged, thought, "huh, wonder who that's affecting." Opened up Age of Empires II....no connection. Go to website for game servers..."Our provider, OVH, is down...." Go figure.

OVH is rather popular in Europe atleast (alongside Hetzner and 1&1).

The Canadian data center is also a very low cost way to serve the US. I'm an OVH fan. They have their quirks, but the pricing is great. You just make sure you compensate for their quirks with backups and DR plans.

Re: OVH Incident in Strasbourg

#98
This affects DNS as well, since domaindiscount24 (a rather large registrar in Germany) happens to host all three of their nameservers with OVH.

Just in case you wonder why your sites don't work, even if you host them somewhere else.

Re: OVH Incident in Strasbourg

#99
post #39

Earlier quoted context omitted.

In this case, I wouldn't be to hard on them. As it appears they lost their main power line, the backup power line and both generators failed and one generator has been restarted now.

What is the point of backup generators if you do not verify that they work every so often? I have a very hard time believing that they actually tested that they worked, because a failure of not one, but both of them.

I wonder if the generators or switching hardware had some kind of stupid IoT thing in them that required a network connection? That'd be one for:

https://twitter.com/internetofshit

Re: OVH Incident in Strasbourg

#100
post #71

Earlier quoted context omitted.

BTW this seems to be a better status page than the one submitted to HN (which is 404ing) http://status.ovh.com/

The status page was down during the outage.

If so, then it's just like Amazon's status page during the AWS outage [1].

Pro-tip: self-hosting status page is maybe not the best idea.

[1] https://twitter.com/awscloud/status/836656664635846656?lang=...

Post reply on HN