Live data from Hacker News

OVH Incident in Strasbourg

status.ovh.com

61–70 of 207 posts

Re: OVH Incident in Strasbourg

#61
post #43
post #4

More info on Twitter from OVH's CEO: https://twitter.com/olesovhcom and on https://twitter.com/ovh_support_en "SBG: ERDF is trying to find out the default. 2 separated 20kV lines are down. We are trying to restart 2 generators A+B for SBG1/SG4. 2 others generators A+B work in SBG2. 1 routing room is in SBG1, the second in SBG2. Both are down. " "An incident is ongoing impacting our network. We are all on the problem.…

ETA is now 30 min for RBX https://twitter.com/olesovhcom/status/928552251458818048

It said the database was corrupted, so yeah, probably routing configuration.

(edit: parent's post was edited after I posted)

Re: OVH Incident in Strasbourg

#62
post #53

Earlier quoted context omitted.

What is the point of backup generators if you do not verify that they work every so often? I have a very hard time believing that they actually tested that they worked, because a failure of not one, but both of them.

To be fair people do test generators on a monthly schedule usually. Problem you find is it’s getting colder now so any problems are amplified suddenly. Might have been entirely tested a couple of weeks ago.

If your generators are not reliable in the cold, the issue is the placement of them. This is something basic to account for.

Re: OVH Incident in Strasbourg

#63
post #53

Earlier quoted context omitted.

To be fair people do test generators on a monthly schedule usually. Problem you find is it’s getting colder now so any problems are amplified suddenly. Might have been entirely tested a couple of weeks ago.

If your generators are not reliable in the cold, the issue is the placement of them. This is something basic to account for.

Yes and no. What the generator vendor says when they sell you the generator and what actually happens 5 years down the line are two different things.

Re: OVH Incident in Strasbourg

#64

Earlier quoted context omitted.

Huh, I was just looking at them too. Contrary to popular opinion, I kinda prefer it when these things happen before I sign up so that in the post-mortem usually whatever architectural failure lead to the outage is corrected and you get a stronger service. Usually.

Wouldn't this require that there be a very limited amount of things that can go wrong? Unless by stronger you mean some kind of a change in mentality over care for the service, but that would probably have a limited lifespan until it's back to normal.

Personally, I like to know how a service fails and recovers.

If a service never fails then I don't know how well they can recover. I don't know anything about their failure mode.

In this case, I'm learning how OVH handles failure modes, how well they handle it, etc.

I can observe how they will treat such things in the future.

Re: OVH Incident in Strasbourg

#66
post #49
post #30

Earlier quoted context omitted.

Spot on. Having one AWS account scares the crap out of me as well. It’s never a good thing if all your eggs are in one basket. My money is on stuff spread across Bytemark, Linode and DigitalOcean with a DR plan involving mostly automatic recovery. AWS doesn’t get a look in as it is extremely costly to port away from anything that isn’t bare metal and pipes.

How do you do fail-over ?

Mirror your data onto another provider continuously (log shipping/rsync), Switch DNS.

Ansible works for this stuff as it allows you to can the task of “quick get me a production environment up on Linode!”

If you can afford some downtime you don’t need a hot standby just roll out everything into new provider and you’re done.

I’ve done this on very large scale environments and small ones and it’s achievable for even small organisations. The killer is avoid anything you can’t run on bare metal servers.

Re: OVH Incident in Strasbourg

#67
post #59

Trending on Twitter with the hashtag #OVHGATE https://twitter.com/hashtag/OVHGATE?src=hash

We selected their three data center EU region precisely because they were three separate data centers, so not happy. This is clearly bad design.

I think we're now going to have to look into multi-provider options. The only way to be solidly up is to be hosted by more than one company at more than one data center.

I've also heard stories of billing nightmares where you get locked out of a cloud provider account, so that's another thing.

Re: OVH Incident in Strasbourg

#68
post #2

Some servers in GRA still appear to work if that's of any help. All data centres offline at once sounds more like an attack than a power failure in one location. According to them, there was a power failure in SBG but I don't see how that should affect routing in data centres several hundred miles away. https://twitter.com/olesovhcom/status/928541667283623936 EDIT: Maybe related to the Cisco issue? https://blogs.cisc…

I have a GRA VPS running fine as well. Can't access the control panel or anything else really, though.

Re: OVH Incident in Strasbourg

#70

wow, yesterday I was playing with their public cloud because considering choosing them. I had some connection problem with my private networking there (deleted it more than once) and opened a ticket. If it was me... sorry, haha. Not good advertisement but it can happen to everyone.

Huh, I was just looking at them too. Contrary to popular opinion, I kinda prefer it when these things happen before I sign up so that in the post-mortem usually whatever architectural failure lead to the outage is corrected and you get a stronger service. Usually.

Not saying this is the case here but that's also sometimes where you spot amateurism and should run away. I remember a host provider a long time ago (15y) who was storing its backups on the same machine as the main data. Guess how I figured out!
Post reply on HN