Live data from Hacker News

Fire declared in OVH SBG2 datacentre building

travaux.ovh.net

501–510 of 613 posts

Re: Fire declared in OVH SBG2 datacentre building

#501

in case there were not posted before, here are pictures of SBG2 in flames taken by the firemen. https://twitter.com/xgarreau/status/1369559995491172354 This puts an image on the sentence "SBG2 is destroyed". Do not expect any recovery from SBG2.

Holy hell. Are these "datacenters" really just shipping containers? That's what it looks like.

Aren't they a bargain-basement provider? You get what you pay for I guess?

Re: Fire declared in OVH SBG2 datacentre building

#503

Earlier quoted context omitted.

There are always uninsurable events and for large enough companies/risks there are also liquidity limits to the size of coverage you can get from the market even for insurable events. As such, it makes sense to make the level of risk you plan to accept (by not being insured against it and not mitigating) a conscious economic decision rather than pretending you've covered everything.

As long as you have outside shareholders you can decide that. If you do you'd be surprised about how they will respond to an attitude like that. After all: you can decide the levels of risk that you personally are comfortable with leading to extinguishing of the business, but a typical shareholder is looking at you to protect their investment and not insuring against a known risk which at some point in time materiali…

In my work life I am a professional investor, so I've been through the debate on insure/prepare or not many times. It's always an economic debate when you get into "very expensive" territory (cheap and easy is different obviously).

The big example of this which springs to mind is business interruption cover - it's ruinously expensive so it's extremely unusual to have the max cover the market might be prepared to offer. It's a pure economic decision.

Re: Fire declared in OVH SBG2 datacentre building

#504

Back in the late 90s, I implemented the first systematic monitoring of WalMart Store's global network, including all of the store routers, hubs (not switches yet!) and 900mhz access points. Did you know that WalMart had some stores in Indonesia? They did until 1998. So when https://en.wikipedia.org/wiki/May_1998_riots_of_Indonesia started happening, we heard some harrowing stories of US employees being abducted, amon…

Thanks for sharing this interesting story. Part of my family immigrated from Indonesia due to those riots, but I was unaware up until today of the details covered by the Wikipedia article you linked. I remember during the 2000s and 2010s that WalMart in the USA earned a reputation for it's inventories primarily consisting of Chinese-made goods. I'm not sure if that reputation goes all the way back to 1998, but it mak…

What the hell, here's another story. The summary to catch your attention: in the early 2000s, I first became aware of WalMart's full scale switch to product sourcing from China by noting some very unusual automated network to site mappings.

Part of what my team (Network Management) did was write code and tools to automate all of the various things that needed to be done with networking gear. A big piece of that was automatically discovering the network. Prior to our auto discovery work, there was no good data source for or inventory of the routers, hubs, switches, cache engines, access points, load balancers, VOIP controllers...you name it.

On the surface, it seems scandalous that we didn't know what was on our own network, but in reality, short of comprehensive and accurate auto discovery, there was no way to keep track of everything, for a number of reasons.

First was the staggering scope: when I left the team, there were 180,000 network devices handling the traffic for tens of millions of end nodes across nearly 5,000 stores, hundreds of distribution centers and hundreds of home office sites/buildings in well over a dozen countries. The main US Home Office in Bentonville, Arkansas was responsible for managing all of this gear, even as many of the international home offices were responsible for buying and scheduling the installation of the same gear.

At any given time, there were a dozen store network equipment rollouts ongoing, where a 'rollout' is having people visit some large percentage of stores intending to make some kind of physical change: installing new hardware, removing old equipment, adding cards to existing gear, etc.

If store 1234 in Lexington, Kentucky (I remember because it was my favorite unofficial 'test' store :) was to get some new switches installed, we would probably not know what day or time the tech to do the work was going to arrive.

ANYway...all that adds up to thousands of people coming in and messing with our physical network, at all hours of the day and night, all over the world, constantly.

Robust and automated discovery of the network was a must, and my team implemented that. The raw network discovery tool was called Drake, named after this guy: https://en.wikipedia.org/wiki/Francis_Drake and the tool that used many automatic and manual rules and heuristics to map the discovered networking devices to logical sites (ie, Store 1234, US) was called Atlas, named after this guy: https://en.wikipedia.org/wiki/Atlas_(mythology)

All of that background aside, the interesting story.

In the late 90s and early 2000s, Drake and Atlas were doing their thing, generally quite well and with only a fairly small amount of care and feeding required. I was snooping around and noticed that a particular site of type International Home Office had grown enormously over the course of a few years. When I looked, it had hundreds of network devices and tens of thousands of nodes. This was around 2001 or 2002, and at that time, I knew that only US Home Office sites should have that many devices, and thought it likely that Atlas had a 'leak'. That is, as Atlas did its recursive site mapping work, sometimes the recursion would expand much further than it should, and incorrectly map things.

After looking at the data, it all seemed fine. So I made some inquiries, and lo and behold, that particular international home office site had indeed been growing explosively.

The site's mapped name was completely unfamiliar to me, at the time at least. You might have heard of it: https://en.wikipedia.org/wiki/Shenzhen

I was seeing fingerprints in our network of WalMart's whole scale switch to sourcing from China.

Re: Fire declared in OVH SBG2 datacentre building

#505

Earlier quoted context omitted.

Newly constructed datacenters in the US tend to be all metal with a full building clean suppression agent. https://www.fike.com/products/ecaro-25-clean-agent-fire-supp... I used to work for a provider whose 2 main datacenters of 8k+ sq ft could pull all oxygen out of the building in 60 seconds.

Data centres I used to work in back in the early 2000s had argonite gas dumps in place (prior to argonite, halon used to be popular but is an ozone depleting gas so was phased out) In the case of a fire, it would dump a lot of argonite gas in and consume a large amount of the oxygen in the room, depriving the fire of fuel. It's also safe and leaves minimal clean-up work afterwards, doesn't harm electronics etc. unlik…

Well, yeah, these normal inert gas fire suppression systems don't do a good job if humans can still breathe. The Novec 1230 based ones can actually be sufficiently effective for typical flammability properties you can cheaply adhere to in a datacenter, but even then you iirc would want to add both that and some extra oxygen, because the nitrogen in the air is much more effective at suffocating humans than at suffocating fire. This stuff is just a really, really heavy gas that's liquid below about body temperature (boils easily though), and the heat capacity of gasses is mostly proportional to their density.

Flames are extinguished by this cooling effect (identical to water in that regard), but humans rely on catalytic processes that aren't affected by the cooling effect.

If you could keep the existing oxygen inside, while adding Novec 1230, humans could continue to breathe while the flames would still be extinguished, but this would require the building/room to be a pressure chamber that holds about half an atmosphere of extra pressure. I'm pretty sure just blowing in some extra oxygen with the Novec 1230 would be far cheaper to do safely and reliably.

I mean, in principle, if you gave re-breathers to the workers and have some airlocks, you could afford to keep that atmosphere permanently, but it'd have to be a bit warm (~30 C I'd guess). Don't worry, the air would be breathable, but long-term it'd probably be unhealthy to breathe in such high concentrations and humans breathing would slightly pollute the atmosphere (CO2 can't stay if it's supposed to remain breathable).

Just to be clear: in an effective argonite extinguishing system you'd have about a minute or two until you pass out and need to be dragged out, ideally get oxygen, get ventilated (no brain, no breathing) and potentially also be resuscitated (the heart stops shortly after your brain from a lack of oxygen, so if you're ventilated fast enough, it never stops and you wake up a few externally-forced breaths later). Having an oxygen bottle to supplement your breaths would fix that problem for as long as it's not empty.

Re: Fire declared in OVH SBG2 datacentre building

#506

The classic "lp0 on fire" error message comes to mind: https://en.wikipedia.org/wiki/Lp0_on_fire Really though, I feel truly awful for anyone affected by this. The post recommends implementing a disaster recovery plan. The truth is that most people don't have one. So, let's use this post to talk about Disaster Recovery Plans! Mine: I have 5 servers at OVH (not at SBG) and they all back up to Amazon S3 or Backblaze B2…

Personally my email hosting is down but thanksfully my web hosting and nextcloud instance were both at GRA2 (Gravelines).

But i have a friend who potentially lost important uni work hosted on his nextcloud instance... On SBG2.

A rough reminder that backups are really important, even if you are just an individual

Re: Fire declared in OVH SBG2 datacentre building

#507

Earlier quoted context omitted.

As long as you have outside shareholders you can decide that. If you do you'd be surprised about how they will respond to an attitude like that. After all: you can decide the levels of risk that you personally are comfortable with leading to extinguishing of the business, but a typical shareholder is looking at you to protect their investment and not insuring against a known risk which at some point in time materiali…

In my work life I am a professional investor, so I've been through the debate on insure/prepare or not many times. It's always an economic debate when you get into "very expensive" territory (cheap and easy is different obviously). The big example of this which springs to mind is business interruption cover - it's ruinously expensive so it's extremely unusual to have the max cover the market might be prepared to offe…

Yes, but it is an informed decision and typically taken at the board level, very few CEO's that are not 100% owners would be comfortable with the decision to leave an existential risk uncovered without full approval of all those involved, which is kind of logical.

Usually you'd have to show your homework (offers from insurance companies proving that it really is unaffordable). I totally get the trade-off, and the fact that if the business could not exist if it was properly insured that plenty of companies will simply take their chances.

We also both know that in case something like that does go wrong everybody will be looking for a scapegoat, so for the CEO's own protection it is quite important to play such things by the book, on the off chance the risk one day does materialize.

Re: Fire declared in OVH SBG2 datacentre building

#508

Earlier quoted context omitted.

Is that a typo? I only see OVH bare metal starting at >$50. How could a provider offer a bare metal server for $5?

It's not a typo. OVH runs Kimsufi, which has bare metal servers for as low as 5$. It is pretty insane.

Thank you. TIL!

Re: Fire declared in OVH SBG2 datacentre building

#509

Earlier quoted context omitted.

Thanks for sharing this interesting story. Part of my family immigrated from Indonesia due to those riots, but I was unaware up until today of the details covered by the Wikipedia article you linked. I remember during the 2000s and 2010s that WalMart in the USA earned a reputation for it's inventories primarily consisting of Chinese-made goods. I'm not sure if that reputation goes all the way back to 1998, but it mak…

What the hell, here's another story. The summary to catch your attention: in the early 2000s, I first became aware of WalMart's full scale switch to product sourcing from China by noting some very unusual automated network to site mappings. Part of what my team (Network Management) did was write code and tools to automate all of the various things that needed to be done with networking gear. A big piece of that was a…

Epic story! Thank you for sharing it. I appreciate the detail you included there.

Re: Fire declared in OVH SBG2 datacentre building

#510

Earlier quoted context omitted.

What the hell, here's another story. The summary to catch your attention: in the early 2000s, I first became aware of WalMart's full scale switch to product sourcing from China by noting some very unusual automated network to site mappings. Part of what my team (Network Management) did was write code and tools to automate all of the various things that needed to be done with networking gear. A big piece of that was a…

Epic story! Thank you for sharing it. I appreciate the detail you included there.

You're quite welcome. In my experience, the right details often make a story far more interesting.
Post reply on HN