Live data from Hacker News

Fire declared in OVH SBG2 datacentre building

travaux.ovh.net

81–90 of 613 posts

Re: Fire declared in OVH SBG2 datacentre building

#81
post #65

I don’t know what ovh is and going to the site point me to a speed test with no information.

French AWS/GCP is my understanding.

They are really more like Hetzner. They have "cloud", but most of the business is dedicated servers. They also operate kimsufi.com and soyoustart.com.

They do have APAC, Canadian, and US data centers as well.

Re: Fire declared in OVH SBG2 datacentre building

#82

Reminder to not only have backups, but also have some periodic OFFLINE backups. If your primary is set up with credentials to automatically transfer a copy to the backup destination over the network, what happens if your primary gets pwned and the access is used to encrypt or delete the backup? Secondly, test doing restores of your backups, and have methods/procedures in place for exactly what a restore looks like.

[deleted]

Re: Fire declared in OVH SBG2 datacentre building

#83
post #67

A status update on the OVH tracker for a different datacenter (LIM-1 / Limburg) says "We are going to intervene in the rack to replace a large number of power supply cables that could have an insulation defect." [0][1] The same type of issue is "planned" in BHS [3] and GRA [2]. Eerie timing: do they possibly suspect some bad cables? [0]: http://travaux.ovh.net/?do=details&id=49016 [1]: http://travaux.ovh.net/?do=deta…

They're waiting an awful long time to do the one at BHS-7 if so: 14 days from now?

Re: Fire declared in OVH SBG2 datacentre building

#84
post #46
post #43

Earlier quoted context omitted.

That's true, but it seems whole of SBG region for OVH is within same disaster radius for one fire... with SBG2 destroyed and SBG1 partly damaged. "The whole site has been isolated, which impacts all our services on SBG1, SBG2, SBG3 and SBG4. " Wonder if those SBGx were advertised as being the same as "Availability Zones" - when other cloud providers ensure zones are distanced enough from each other (~1km at least) to…

Thats a fair point. If OVH does market them as AZs then it's disingenuous and liable to suits IMO.

No, it isn't, as there's no clear cut definition of what an availability zone is.

Re: Fire declared in OVH SBG2 datacentre building

#85

The classic "lp0 on fire" error message comes to mind: https://en.wikipedia.org/wiki/Lp0_on_fire Really though, I feel truly awful for anyone affected by this. The post recommends implementing a disaster recovery plan. The truth is that most people don't have one. So, let's use this post to talk about Disaster Recovery Plans! Mine: I have 5 servers at OVH (not at SBG) and they all back up to Amazon S3 or Backblaze B2…

If you are a corporate entity of some kind, the final layer of your plan should always be "Go bankrupt". You can't successfully recover from every possible disaster and you shouldn't try to. In the event of a sufficiently unlikely event, your business fails and every penny spent attempting the impossible will be wasted, move on and let professional administrators salvage what they can for your creditors. Lots of peop…

IMHO, the part they had no plan for was being unable to just require their employees to come in anyway...

Re: Fire declared in OVH SBG2 datacentre building

#86
post #17

Unfortunately A lot of people are going to find out the hard way today why AWS/GCP/Big Expensive Cloud is so expensive (Hint: they have redundancy and failover procedures which drive up costs). Keep in mind I’m talking not of “downtime” but of actual data loss which might affect business continuity. This is really tragic. I’m hoping they have some kind of multi regional backup/replication and not just multi zones (al…

I encourage you to have a look at the operating income that AWS rakes in. Sure, the amount of expertise, redundancy and breadth of service offerings they provide is worth a markup, but they are also significantly more expensive than they need to be. Thanks to being the leader in an oligopoly, and due to patterns like making network egress unjustifiably expensive to keep you (/your data) from leaving.

I think the question here, then is of subjective value.

AWS may charge more for egress, but that’s not high enough for it to be a concern for most clients.

A bigger, independent concern is probably that there should be sufficient redundancy, backups and such that allows for business continuity. (Note again that I’m not saying that all companies make full use of these features, but those that care for such things do. Additionally, I’ve honestly never heard of an AWS DC burning down. Either it doesn’t happen frequently or it doesn’t have enough of an effect on regular customers, both of situations are equivalent for my case).

Most businesses choose to prioritize the second aspect. Even if they have to pay extra for egress sometimes, it’s just not big enough of a concern as compared to businesses continuity.

Re: Fire declared in OVH SBG2 datacentre building

#87
(site a)---[replicate local LUNs/shares to remote storage arrays]--->(site b), (site a)---[replicates local VMs to remote HCI]--->(site b), (site a)---[local backups to local data archive]--->(site a), (site a)---[local data archive replicates to remote data archive]--->(site b), (site b)---[remote data archive replicates to remote air gapped data archive]--->(site b), (site a)---[replicates to cold storage on aws/gcp/azure]--->(site c), (site c)---[replicate to another geo site on cloud]--->(site d)

scenario 1: site a is down plan: recover to site b by most convinent means

scenario 2: site b is down plan: restore services, operate without redundancy out of site a

scenario 3: site c is down plan: restore services, catch up later. continue operating out of site a

scenario 4: site b and c down plan: restore services, operate without redundancy out of site a

scenario 5: site a and b down plan: cross fingers, restore to new site from cold storage on expensive cloud VM instances

scenario 6: data archive corrupted ransomware plan: restore from air gapped data archive, hope ransomware was identified within 90 days

scenario 7: site b and c down, then site a down plan: quit

scenario 8: staff hates job and all quit plan: outsource

scenario 9: and so on...

Re: Fire declared in OVH SBG2 datacentre building

#88
post #17

Unfortunately A lot of people are going to find out the hard way today why AWS/GCP/Big Expensive Cloud is so expensive (Hint: they have redundancy and failover procedures which drive up costs). Keep in mind I’m talking not of “downtime” but of actual data loss which might affect business continuity. This is really tragic. I’m hoping they have some kind of multi regional backup/replication and not just multi zones (al…

“Big cloud” has had fires take out clusters, and somehow they manage to keep it out of the news. In spite of the redundancy and failover procedures, keeping your data centers running when one of the clusters was recently *on fire* is something that is often only possible due to heroic efforts. When I say “heroic efforts”, that’s in contrast to “ordinary error recovery and failover”, which is the way you’d want to han…

You know, those are excellent observations. But they don’t change the decision calculus in this case. Using bigger cloud providers doesn’t eliminate all risk, it just creates a different kind of risk.

What we call “progress” in humanity is just putting our best efforts into reducing or eliminating the problems we know how to solve without realizing the problems they may create further down the line. The only way to know for sure is to try it, see how it goes, and then re-evaluate later.

California had issues with many forest fires. They put out all fires. Turns out, that solution creates a bigger problem down the line with humongous uncontrollable fires which would not have happened if the smaller fires had not been put out so frequently. Oops.

Re: Fire declared in OVH SBG2 datacentre building

#89
post #14

The classic "lp0 on fire" error message comes to mind: https://en.wikipedia.org/wiki/Lp0_on_fire Really though, I feel truly awful for anyone affected by this. The post recommends implementing a disaster recovery plan. The truth is that most people don't have one. So, let's use this post to talk about Disaster Recovery Plans! Mine: I have 5 servers at OVH (not at SBG) and they all back up to Amazon S3 or Backblaze B2…

One of my backup servers used to be in the same datacenter as the primary server. I only recently moved it to a different host. It's still in the same city, though, so I'm considering other options. I'm not a big fan of just-make-a-tarball-of-everything-and-upload-it-to-the-cloud backup methodology, I prefer something a bit more incremental. But with Backblaze B2 being so cheap, I might as well just upload tarballs t…

> I'm not a big fan of just-make-a-tarball-of-everything-and-upload-it-to-the-cloud backup methodology, I prefer something a bit more incremental.

pretty much a textbook use-case for zfs with some kind of snapshot-rolling utility. Snap every hour, send backups once a day, prune your backups according to some timetable. Transfer as incrementals against the previous stored snapshot. Plus you get great data integrity checking on top of that.

"but linus said..."

Re: Fire declared in OVH SBG2 datacentre building

#90
I can't see anything about a fire suppression system mentioned? Doesn't OVH have one, except for colocation datacenters?

A fire detection system using eg. lasers and Inergen(or Argonite) for putting the fire out is commonly used in datacenters. The gas fills the room and reduces the amount of oxygen in the room so most fires are put out within a minute.

The cool thing is that the gas is designed to be used in rooms with people, so that is can be triggered any time. It is however quite loud, and some setups have been known to be too loud, even destroying harddrives.

Post reply on HN