Live data from Hacker News

OVH Incident in Strasbourg

status.ovh.com

41–50 of 207 posts

Re: OVH Incident in Strasbourg

#41

Someone with access might wish to update the title of this post, because all OVH datacenters are definitely not down.

But no one knows which DCs and services are down. They lost their internal network and have no idea themselves.

Re: OVH Incident in Strasbourg

#43
post #4

More info on Twitter from OVH's CEO: https://twitter.com/olesovhcom and on https://twitter.com/ovh_support_en "SBG: ERDF is trying to find out the default. 2 separated 20kV lines are down. We are trying to restart 2 generators A+B for SBG1/SG4. 2 others generators A+B work in SBG2. 1 routing room is in SBG1, the second in SBG2. Both are down. " "An incident is ongoing impacting our network. We are all on the problem.…

ETA is now 30 min for RBX

https://twitter.com/olesovhcom/status/928552251458818048

Re: OVH Incident in Strasbourg

#44

A website that I host on ovh is up: https://sudokugarden.de/ ovh.com looks down for me too. You can check it's hosted by OVH: $ whois $(dig sudokugarden.de +short)

It's in the Gravelines datacenter, which is not actually down contrary to what the initial reports said (only Strasbourg and Roubaix are, and for two different reasons).

Re: OVH Incident in Strasbourg

#45
post #26

Earlier quoted context omitted.

That seems a little uncharitable.

If the SBG issue really triggered the outage for the whole network I find it hard to believe that no one saw that problem beforehand. They probably thought that this was too unlikely to happen or that there are other failovers but never tested them properly. No expert on the field but that's the first time I can remember that a provider of that size loses connection to most of their data centres at once. That can hap…

> No expert on the field but that's the first time I can remember that a provider of that size loses connection to most of their data centres at once.

It happened to GCE last year[1], though it only lasted 18 minutes.

[1]: https://status.cloud.google.com/incident/compute/16007?post-...

Re: OVH Incident in Strasbourg

#46
post #39
post #26

Earlier quoted context omitted.

If the SBG issue really triggered the outage for the whole network I find it hard to believe that no one saw that problem beforehand. They probably thought that this was too unlikely to happen or that there are other failovers but never tested them properly. No expert on the field but that's the first time I can remember that a provider of that size loses connection to most of their data centres at once. That can hap…

In this case, I wouldn't be to hard on them. As it appears they lost their main power line, the backup power line and both generators failed and one generator has been restarted now.

So what? Losing main power is a standard case for any DC. That's why you have generators. Even a generator failure is nothing out of the ordinary. But that no generators in a DC work kind of indicates that they don't test them as often as you would expect.

They just announced that they want to be a "hypercloud" provider on the scale of AWS and Google Cloud. I really hope that a power failure in Virginia couldn't bring down all of AWS.

Re: OVH Incident in Strasbourg

#47
post #16

I moved away from OVH after I paid 3 months advance (~$300) for a server which burned down after 1 1/2 months. They did not issue any refunds (data, blood, sweat and tears were lost that day). I have been an OVH customer for 12 years. Today, I'm glad to have moved away all my production environments as well.

My experience with OVH VPSes is that they are cheap and reasonably reliable. But when one starts failing support never helped me to get it fixed.

Just fire up another one, install everything on there, switch off the old one and forget about it. This means never paying for more than one month in advance.

Re: OVH Incident in Strasbourg

#48
post #39
post #26

Earlier quoted context omitted.

If the SBG issue really triggered the outage for the whole network I find it hard to believe that no one saw that problem beforehand. They probably thought that this was too unlikely to happen or that there are other failovers but never tested them properly. No expert on the field but that's the first time I can remember that a provider of that size loses connection to most of their data centres at once. That can hap…

In this case, I wouldn't be to hard on them. As it appears they lost their main power line, the backup power line and both generators failed and one generator has been restarted now.

What is the point of backup generators if you do not verify that they work every so often? I have a very hard time believing that they actually tested that they worked, because a failure of not one, but both of them.

Re: OVH Incident in Strasbourg

#49
post #30
post #19

Earlier quoted context omitted.

I'm still at OVH (support is reasonable and prices cheap) but would never trust one provider with all my infrastructure. In the end it always turns out that there is a single point of failure and if it's just the billing department. Using two providers protects you from that and if you chose some with good peering and free traffic, keeping both in sync is relatively easy. That's just the problem with services as AWS,…

Spot on. Having one AWS account scares the crap out of me as well. It’s never a good thing if all your eggs are in one basket. My money is on stuff spread across Bytemark, Linode and DigitalOcean with a DR plan involving mostly automatic recovery. AWS doesn’t get a look in as it is extremely costly to port away from anything that isn’t bare metal and pipes.

How do you do fail-over ?
Post reply on HN