Live data from Hacker News

Postmortem of the failure of one hosting storage unit on Jan. 8, 2020

news.gandi.net

51–60 of 83 posts

Re: Postmortem of the failure of one hosting storage unit on Jan. 8, 2020

#51

And now their postmortem blog post is down with a 503 error. Doesn’t exactly fill me with confidence about their abilities.

Their handling of the issue on Twitter was enough for to decide to move my domain names away from them when their renewal is due.

The last I looked at them for either hosting or domain, they had a provision in their TOS that basically said they could terminate my account at any time if I did anything they felt was morally wrong. I emailed them and asked if what I was reading was true and they confirmed it. I never looked at them again.

They probably thought their no-BS morally-right stance was supposed to comfort me, and it's not like I host anything that would likely meet that criteria. But who is to say a blog post has cursing in it that they decide, at their sole discretion, to be bad? Or speak up on abortion in a way they don't agree with? Or any of the other morally-charged topics out there? I'm not hosting with the thought police and it's always made me wonder how others felt comfortable with them.

Re: Postmortem of the failure of one hosting storage unit on Jan. 8, 2020

#52
post #29
post #13

Just bad luck. A different story: I had a cheap dedicated host in Atlanta. Their failure was epic. You get what you pay for. An old highrise is filled with ten thousands of old second hand server blades, floors by floors of equpiment prolifically producing waste heat. A sure recipe for disaster? Sure! A wrongly installed fuse at one phase in the building made one phase burn out too early. I saw a picture of an archeo…

Not bad luck, bad design. Same for Gandi, and for the situation you just described. Drive failures and HVAC failures due to bad power are not "black swawn" events. These are very common problems for DCs, and a good design takes these problems into account. However, a "bad" design is cheap, and hopefully the savings is passed to the customer.

[deleted]

Re: Postmortem of the failure of one hosting storage unit on Jan. 8, 2020

#53

Earlier quoted context omitted.

No offsite backups? Backblaze B2 is exceedingly cheap for example.

The animation studio I worked for had almost a petabyte of data. It may be cheap to buy the storage but transferring is costly. It's very easy to saturate a MPLS circuit with data, even rSync on a 10Gbit internal connection takes a long while. Really Gandi should of had backups from day one. If your hosting data you should always have backups ready and tested on day one.

I’d be very curious if you evaluated the post-MPLS guys like Megaport for connectivity.

Re: Postmortem of the failure of one hosting storage unit on Jan. 8, 2020

#54
post #22

Earlier quoted context omitted.

For those that missed it, an employee handling their twitter account offered nudes as compensation for the outage.

No they didn't, they were referencing the "shame" scene, and even included a gif, from Game of Thrones. In context, them saying "who do you want to see naked" is obviously not offering to send nudes, but who do you want to see punished. The worst part was the CEO getting on twitter and mocking the person who had originally complained about the outage on twitter.

Anyone not familiar with GoT could certainly have easily misinterpreted that though.

Re: Postmortem of the failure of one hosting storage unit on Jan. 8, 2020

#55
post #17

They don't say they're sorry, because they're not. Instead they minimize their actions by: 1) stating how few customers customers were affected, 2) how it's not really their fault because it was a hardware error, 3) it's not really their fault because they had already planned to upgrade the server, 4) it's not really their fault the restore procedure took so long because they had to make backups first, 5) the restore…

>They don't say they're sorry,

>We’re very sorry for this truly unfortunate incident and we offer our sincere apologies to anyone impacted.

https://news.gandi.net/en/2020/01/major-incident-on-our-host... (linked in the Postmortem)

Re: Postmortem of the failure of one hosting storage unit on Jan. 8, 2020

#56

Earlier quoted context omitted.

No data was lost though, true? I don't know if I expect a postmortem to say "sorry", and I think you are being needlessly harsh. But I agree this level of service doesn't seem up to current best in class. Like Amazon etc. (Which of course still have unexpeted outages very occasionally, although a 5 day time to recovery would certainly be... unusual). But this partially shows how much expectations/standards have raise…

Honestly, I too was looking for a straightforward 'We screwed up, sorry.' I wouldn't care nearly as much if they'd just had 5 days without snapshots. But the way they poorly handled support deserves to be addressed in a postmortem. See this thread for the support at the time: https://twitter.com/andreaganduglia/status/12152827193300664...

> Honestly, I too was looking for a straightforward 'We screwed up, sorry.'

They do so a bit here: https://news.gandi.net/en/2020/01/major-incident-on-our-host...

>We’re very sorry for this truly unfortunate incident and we offer our sincere apologies to anyone impacted.

Re: Postmortem of the failure of one hosting storage unit on Jan. 8, 2020

#57
post #51

Earlier quoted context omitted.

Their handling of the issue on Twitter was enough for to decide to move my domain names away from them when their renewal is due.

The last I looked at them for either hosting or domain, they had a provision in their TOS that basically said they could terminate my account at any time if I did anything they felt was morally wrong. I emailed them and asked if what I was reading was true and they confirmed it. I never looked at them again. They probably thought their no-BS morally-right stance was supposed to comfort me, and it's not like I host an…

> they had a provision in their TOS that basically said they could terminate my account at any time if I did anything they felt was morally wrong.

Don't most modern service providers have a clause like that?

Re: Postmortem of the failure of one hosting storage unit on Jan. 8, 2020

#58

And now their postmortem blog post is down with a 503 error. Doesn’t exactly fill me with confidence about their abilities.

Their handling of the issue on Twitter was enough for to decide to move my domain names away from them when their renewal is due.

Any recommendations? Preferred within the EU.

Re: Postmortem of the failure of one hosting storage unit on Jan. 8, 2020

#59
It's interesting -- they ID the actual cause of the problem up top, and then just zip right past it. The problem wasn't the hardware failure, or the lack of backups, it was that customers expected them to have backups.

Gandi goes into some detail on the recovery process and on ways to fix the issue in the future. But, apart from some hand-waving, they don't have any specifics about how they'll communicate expectations better with their customers in the future.

Imagine the counterfactual: Gandi's docs clearly communicate "this service has no backups, you can take a snapshot through this api, you're on your own." Of course customers with data loss would've complained, but, at the end of the day, the message from both Gandi and the community would've been "well, next time buy a service with backups?" Yet there's no explicit plan to improve documentation.

Re: Postmortem of the failure of one hosting storage unit on Jan. 8, 2020

#60
post #32

> We think it may be due to a hardware problem linked to the server RAM. Are they using ECC RAM?

Sounds like they didn't and the metadata logs got corrupted..

This does also mean that other data would be corrupted too, running ZFS without ECC RAM is frequently warned against.

Post reply on HN