Live data from Hacker News

OVH Incident in Strasbourg

status.ovh.com

51–60 of 207 posts

Re: OVH Incident in Strasbourg

#51

wow, yesterday I was playing with their public cloud because considering choosing them. I had some connection problem with my private networking there (deleted it more than once) and opened a ticket. If it was me... sorry, haha. Not good advertisement but it can happen to everyone.

Huh, I was just looking at them too. Contrary to popular opinion, I kinda prefer it when these things happen before I sign up so that in the post-mortem usually whatever architectural failure lead to the outage is corrected and you get a stronger service.

Usually.

Re: OVH Incident in Strasbourg

#52
post #18

I imagine Mr Good Guy at OVH telling some others: "guys we have a single point of failure in our architecture with SBG, maybe we should... - naaah it's fine, we do not have time nor resources" Then shit happens. edit: I have no idea what is happening exactly, but OVH being what it is, it seems extremely weird that all datacenters "can" get down at the same time, and it looks like a serious architecture problem to me…

its OVH: The Hardware is good DDoS protection is good The Prices are high

but

support/administration does not work well, i have a lot of really weird story's with them, from them plugging in a keyboard in our server to reboot it (without any reason) to taking down a server for a requested maintenance only to notice after 4 hours of downtime that they did not ask their bosses if they were allowed to even perform the maintenance requested (and then not getting permission to do so after another 2 hours ..)

For me it feels like there are some really deep issues somewhere in the whole administration that make incidents like this no real surprise

Problem is most other providers dont work any better, so ...

Everyone makes mistakes, let's just hope they learn from it.

Re: OVH Incident in Strasbourg

#53
post #39

Earlier quoted context omitted.

In this case, I wouldn't be to hard on them. As it appears they lost their main power line, the backup power line and both generators failed and one generator has been restarted now.

What is the point of backup generators if you do not verify that they work every so often? I have a very hard time believing that they actually tested that they worked, because a failure of not one, but both of them.

To be fair people do test generators on a monthly schedule usually. Problem you find is it’s getting colder now so any problems are amplified suddenly. Might have been entirely tested a couple of weeks ago.

Re: OVH Incident in Strasbourg

#54
post #45
post #26

Earlier quoted context omitted.

If the SBG issue really triggered the outage for the whole network I find it hard to believe that no one saw that problem beforehand. They probably thought that this was too unlikely to happen or that there are other failovers but never tested them properly. No expert on the field but that's the first time I can remember that a provider of that size loses connection to most of their data centres at once. That can hap…

> No expert on the field but that's the first time I can remember that a provider of that size loses connection to most of their data centres at once. It happened to GCE last year[1], though it only lasted 18 minutes. [1]: https://status.cloud.google.com/incident/compute/16007?post-...

But that's one product, not the whole DC. As I understand the post, other services worked correctly during that time.

Re: OVH Incident in Strasbourg

#55
post #46
post #39

Earlier quoted context omitted.

In this case, I wouldn't be to hard on them. As it appears they lost their main power line, the backup power line and both generators failed and one generator has been restarted now.

So what? Losing main power is a standard case for any DC. That's why you have generators. Even a generator failure is nothing out of the ordinary. But that no generators in a DC work kind of indicates that they don't test them as often as you would expect. They just announced that they want to be a "hypercloud" provider on the scale of AWS and Google Cloud. I really hope that a power failure in Virginia couldn't brin…

The generators did work but they failed. Both of them.

I can not imagine they weren't tested.

But even the most rigorous testing can never reduce the total failure risk to 0. It seems OVH just got very very unlucky.

Re: OVH Incident in Strasbourg

#56
post #46
post #39

Earlier quoted context omitted.

In this case, I wouldn't be to hard on them. As it appears they lost their main power line, the backup power line and both generators failed and one generator has been restarted now.

So what? Losing main power is a standard case for any DC. That's why you have generators. Even a generator failure is nothing out of the ordinary. But that no generators in a DC work kind of indicates that they don't test them as often as you would expect. They just announced that they want to be a "hypercloud" provider on the scale of AWS and Google Cloud. I really hope that a power failure in Virginia couldn't brin…

When I've seen things like this before, it's often been the switchover hardware that fails, not the generator as such. it's much harder to test that as you don't want to tell your customers "sorry your sever went down, we were just testing if the switchover worked and it didn't"

Re: OVH Incident in Strasbourg

#57
post #41

Someone with access might wish to update the title of this post, because all OVH datacenters are definitely not down.

But no one knows which DCs and services are down. They lost their internal network and have no idea themselves.

I'm pretty sure they know exactly which services are down.

Even if they didn't, clearly many services are up and running normally, so saying "all datacenters are down" is just a lie.

Re: OVH Incident in Strasbourg

#60

wow, yesterday I was playing with their public cloud because considering choosing them. I had some connection problem with my private networking there (deleted it more than once) and opened a ticket. If it was me... sorry, haha. Not good advertisement but it can happen to everyone.

Huh, I was just looking at them too. Contrary to popular opinion, I kinda prefer it when these things happen before I sign up so that in the post-mortem usually whatever architectural failure lead to the outage is corrected and you get a stronger service. Usually.

Wouldn't this require that there be a very limited amount of things that can go wrong?

Unless by stronger you mean some kind of a change in mentality over care for the service, but that would probably have a limited lifespan until it's back to normal.

Post reply on HN