And now their postmortem blog post is down with a 503 error. Doesn’t exactly fill me with confidence about their abilities.
Their handling of the issue on Twitter was enough for to decide to move my domain names away from them when their renewal is due.
Postmortem of the failure of one hosting storage unit on Jan. 8, 2020
21–30 of 83 posts
Re: Postmortem of the failure of one hosting storage unit on Jan. 8, 2020
#22Earlier quoted context omitted.
Their handling of the issue on Twitter was enough for to decide to move my domain names away from them when their renewal is due.
For those that missed it, an employee handling their twitter account offered nudes as compensation for the outage.
Re: Postmortem of the failure of one hosting storage unit on Jan. 8, 2020
#23They don't say they're sorry, because they're not. Instead they minimize their actions by: 1) stating how few customers customers were affected, 2) how it's not really their fault because it was a hardware error, 3) it's not really their fault because they had already planned to upgrade the server, 4) it's not really their fault the restore procedure took so long because they had to make backups first, 5) the restore…
I don't know if I expect a postmortem to say "sorry", and I think you are being needlessly harsh. But I agree this level of service doesn't seem up to current best in class. Like Amazon etc. (Which of course still have unexpeted outages very occasionally, although a 5 day time to recovery would certainly be... unusual).
But this partially shows how much expectations/standards have raised in the past few-10 years. When an unacceptable not up to par level of reliably still involves no data loss, we're doing pretty good. And I think "don't trust Gandi with anything you care about" is probably an exagerated response. But yes, they don't seem to be providing mega-cloud-service-provider level of service.
Re: Postmortem of the failure of one hosting storage unit on Jan. 8, 2020
#24They don't say they're sorry, because they're not. Instead they minimize their actions by: 1) stating how few customers customers were affected, 2) how it's not really their fault because it was a hardware error, 3) it's not really their fault because they had already planned to upgrade the server, 4) it's not really their fault the restore procedure took so long because they had to make backups first, 5) the restore…
No data was lost though, true? I don't know if I expect a postmortem to say "sorry", and I think you are being needlessly harsh. But I agree this level of service doesn't seem up to current best in class. Like Amazon etc. (Which of course still have unexpeted outages very occasionally, although a 5 day time to recovery would certainly be... unusual). But this partially shows how much expectations/standards have raise…
See this thread for the support at the time:
https://twitter.com/andreaganduglia/status/12152827193300664...
Re: Postmortem of the failure of one hosting storage unit on Jan. 8, 2020
#25Earlier quoted context omitted.
I don't have hundreds customers, but I handle hundreds of TiB of data (for a science lab). Issues occur from time to time, and I can assure that these times are very stressful. I am grateful to rely on ZFS, because yet I have never lost any data from people (datasets are often around 10TiB).
No offsite backups? Backblaze B2 is exceedingly cheap for example.
Really Gandi should of had backups from day one. If your hosting data you should always have backups ready and tested on day one.
Re: Postmortem of the failure of one hosting storage unit on Jan. 8, 2020
#26Earlier quoted context omitted.
For those that missed it, an employee handling their twitter account offered nudes as compensation for the outage.
No they didn't, they were referencing the "shame" scene, and even included a gif, from Game of Thrones. In context, them saying "who do you want to see naked" is obviously not offering to send nudes, but who do you want to see punished. The worst part was the CEO getting on twitter and mocking the person who had originally complained about the outage on twitter.
https://twitter.com/andreaganduglia/status/12152827193300664...
Re: Postmortem of the failure of one hosting storage unit on Jan. 8, 2020
#27Am I reading this right? This works out to just ~3.8TiB.
So much drama over, basically, one HDD worth of data?
Re: Postmortem of the failure of one hosting storage unit on Jan. 8, 2020
#28Earlier quoted context omitted.
No data was lost though, true? I don't know if I expect a postmortem to say "sorry", and I think you are being needlessly harsh. But I agree this level of service doesn't seem up to current best in class. Like Amazon etc. (Which of course still have unexpeted outages very occasionally, although a 5 day time to recovery would certainly be... unusual). But this partially shows how much expectations/standards have raise…
Honestly, I too was looking for a straightforward 'We screwed up, sorry.' I wouldn't care nearly as much if they'd just had 5 days without snapshots. But the way they poorly handled support deserves to be addressed in a postmortem. See this thread for the support at the time: https://twitter.com/andreaganduglia/status/12152827193300664...
I still think the fact that there was no data loss, and we're still on the edge of calling it unacceptable incompetence, is worth noting, as to how far our expectations and standards have come. Which is good of course.
Re: Postmortem of the failure of one hosting storage unit on Jan. 8, 2020
#29Just bad luck. A different story: I had a cheap dedicated host in Atlanta. Their failure was epic. You get what you pay for. An old highrise is filled with ten thousands of old second hand server blades, floors by floors of equpiment prolifically producing waste heat. A sure recipe for disaster? Sure! A wrongly installed fuse at one phase in the building made one phase burn out too early. I saw a picture of an archeo…
Drive failures and HVAC failures due to bad power are not "black swawn" events. These are very common problems for DCs, and a good design takes these problems into account.
However, a "bad" design is cheap, and hopefully the savings is passed to the customer.
Re: Postmortem of the failure of one hosting storage unit on Jan. 8, 2020
#30Earlier quoted context omitted.
Honestly, I too was looking for a straightforward 'We screwed up, sorry.' I wouldn't care nearly as much if they'd just had 5 days without snapshots. But the way they poorly handled support deserves to be addressed in a postmortem. See this thread for the support at the time: https://twitter.com/andreaganduglia/status/12152827193300664...
Fair enough. That probably is bad marketing if nothing else, and maybe something else. what you're saying about the major failure being in support/customer-management, even more than the technical issue, seems potentially reasonable. (I am not a Gandi customer, so it's not personal for me). I still think the fact that there was no data loss, and we're still on the edge of calling it unacceptable incompetence, is wort…
> Hi Andrea. It is confirmed we have lost data and we are terribly sorry for that. However, please note that what happenend[sp] could happen to any web host.
Customers that were forced to migrate to a different webhost had to restore from whatever backups they had, and they lost data for sure. Even if Gandi ultimately recovered everything (and it's not completely clear if they did) at that point the customer data/databases have already been forked so it's too late.