Live data from Hacker News

OVH Incident in Strasbourg

status.ovh.com

21–30 of 207 posts

Re: OVH Incident in Strasbourg

#21
wow, yesterday I was playing with their public cloud because considering choosing them. I had some connection problem with my private networking there (deleted it more than once) and opened a ticket. If it was me... sorry, haha. Not good advertisement but it can happen to everyone.

Re: OVH Incident in Strasbourg

#22

It started with all our SBG servers going down simultaneously. Approximately 1h later all our RBX servers went down as well including the OVH status page and all other OVH web applications. Either their SBG and RBX data centers are somehow connected or those are indeed two independent incidents.

SBG was a power failure, followed by a generator failure. Not sure if they set up their network infrastructure in a way that this could spread but I find it hard to imagine that the outage in SBG triggered RBX going completely black. Unless of course they store their configuration files in SBG.

Re: OVH Incident in Strasbourg

#23
post #12
post #8

Earlier quoted context omitted.

It seems more likely that their data centers aren't quite as isolated as they thought they'd be. The outage also appears to be limited to their locations in Europe.

I always wondered why they promote having 6 datacentres in Roubaix when google maps shows that they're all within 50m. Can't be too much redundancy there.

It's their original implantation. The 6 DCs in Roubaix are probably there for more storage capacity, not for redundancy.

Re: OVH Incident in Strasbourg

#24

wow, yesterday I was playing with their public cloud because considering choosing them. I had some connection problem with my private networking there (deleted it more than once) and opened a ticket. If it was me... sorry, haha. Not good advertisement but it can happen to everyone.

I signed up with ovh.com.au last month and testing stuff on it. Down too.

Re: OVH Incident in Strasbourg

#26
post #18

I imagine Mr Good Guy at OVH telling some others: "guys we have a single point of failure in our architecture with SBG, maybe we should... - naaah it's fine, we do not have time nor resources" Then shit happens. edit: I have no idea what is happening exactly, but OVH being what it is, it seems extremely weird that all datacenters "can" get down at the same time, and it looks like a serious architecture problem to me…

That seems a little uncharitable.

If the SBG issue really triggered the outage for the whole network I find it hard to believe that no one saw that problem beforehand. They probably thought that this was too unlikely to happen or that there are other failovers but never tested them properly.

No expert on the field but that's the first time I can remember that a provider of that size loses connection to most of their data centres at once. That can happen with one product (eg S3 failure) but datacentre switches should work even if the rest is on fire.

Re: OVH Incident in Strasbourg

#28
post #2

Some servers in GRA still appear to work if that's of any help. All data centres offline at once sounds more like an attack than a power failure in one location. According to them, there was a power failure in SBG but I don't see how that should affect routing in data centres several hundred miles away. https://twitter.com/olesovhcom/status/928541667283623936 EDIT: Maybe related to the Cisco issue? https://blogs.cisc…

> According to them, there was a power failure in SBG but I don't see how that should affect routing in data centres several hundred miles away. According to them, they do their routing in SBG, so it's plausible that it could lead to all of their network being down.

How is that even a thing? Don't you have separate routing for every switch? Otherwise you don't have redundancy? I'd even expect data centres to use different hardware for networks to avoid having a single point of failure.

Re: OVH Incident in Strasbourg

#30
post #19
post #16

I moved away from OVH after I paid 3 months advance (~$300) for a server which burned down after 1 1/2 months. They did not issue any refunds (data, blood, sweat and tears were lost that day). I have been an OVH customer for 12 years. Today, I'm glad to have moved away all my production environments as well.

I'm still at OVH (support is reasonable and prices cheap) but would never trust one provider with all my infrastructure. In the end it always turns out that there is a single point of failure and if it's just the billing department. Using two providers protects you from that and if you chose some with good peering and free traffic, keeping both in sync is relatively easy. That's just the problem with services as AWS,…

Spot on.

Having one AWS account scares the crap out of me as well. It’s never a good thing if all your eggs are in one basket.

My money is on stuff spread across Bytemark, Linode and DigitalOcean with a DR plan involving mostly automatic recovery.

AWS doesn’t get a look in as it is extremely costly to port away from anything that isn’t bare metal and pipes.

Post reply on HN