Live data from Hacker News

Postmortem of the failure of one hosting storage unit on Jan. 8, 2020

news.gandi.net

31–40 of 83 posts

Re: Postmortem of the failure of one hosting storage unit on Jan. 8, 2020

#31

This is basically the stuff of nightmares. You can't really fault them for the zfs version being so old the feature they needed wasn't yet implemented, because the machine was literally part of the last batch to be upgraded. The root cause is just some random hardware failure that can't be anticipated. Just bad luck. Beyond radically changing how their core infrastructure works, doesn't seem like there was a lot they…

There was one, obvious thing they could have done. Backups. ZFS even makes it easy. They state they have triple redundancy on their servers. Pull 1/2 of the drives and have backups. ZFS supports streaming snapshots (maybe not on this particularly old system). It sounds like they have multiple ZFS servers per datacenter, so given 2 servers, they could use 50% of the storage on each nodes as a backup of the other node.…

They do mention in the postmortem that they explicitly do not provide backups, and say so on their product page, but perhaps that it could be better communicated to customers.

Designing a really robust system to failures like this is a very difficult problem. You can see this in the complexity of systems like S3 and Google's Colossus[1]. Colossus in particular is probably one of Google's single greatest competitive advantages, especially considering none of it is open sourced[2].

Comparing these guys to AWS/S3 is perhaps not entirely fair given the assumption that they have very different levels of resources. For a medium size shop and the constraints they've defined, I think this is a fair outcome of the situation. I agree though in that it could have been mitigated by making the decision to actually store backups.

[1]https://www.wired.com/2012/07/google-colossus/

[2]https://cloud.google.com/files/storage_architecture_and_chal...

Re: Postmortem of the failure of one hosting storage unit on Jan. 8, 2020

#33
post #17

They don't say they're sorry, because they're not. Instead they minimize their actions by: 1) stating how few customers customers were affected, 2) how it's not really their fault because it was a hardware error, 3) it's not really their fault because they had already planned to upgrade the server, 4) it's not really their fault the restore procedure took so long because they had to make backups first, 5) the restore…

Well, it's a technical post-portem, not a love letter. I'm not affiliated to Gandi in any way, but I find the finger-pointing a bit too pedantic.

> The take-away here is clear: don't trust Gandi with anything you care about.

The take-away is not this one. Its: backup anything you care about.

Re: Postmortem of the failure of one hosting storage unit on Jan. 8, 2020

#34
post #5

Earlier quoted context omitted.

I don't have hundreds customers, but I handle hundreds of TiB of data (for a science lab). Issues occur from time to time, and I can assure that these times are very stressful. I am grateful to rely on ZFS, because yet I have never lost any data from people (datasets are often around 10TiB).

No offsite backups? Backblaze B2 is exceedingly cheap for example.

My lab generates similar-sized data sets and the transfer, more than the at-rest storage, is tough.

If our internet (or Box's datacenter) were slow, we could easily collect data faster than we could send it to our collaborators.

Re: Postmortem of the failure of one hosting storage unit on Jan. 8, 2020

#36
post #21

Earlier quoted context omitted.

Their handling of the issue on Twitter was enough for to decide to move my domain names away from them when their renewal is due.

Just a heads up, and forgive me if this is obvious, but you can move right away, the expire date will be the same anyway. You certainly don't want a failed migration too close to the expiration date.

yes, good point!

Re: Postmortem of the failure of one hosting storage unit on Jan. 8, 2020

#37
post #17

They don't say they're sorry, because they're not. Instead they minimize their actions by: 1) stating how few customers customers were affected, 2) how it's not really their fault because it was a hardware error, 3) it's not really their fault because they had already planned to upgrade the server, 4) it's not really their fault the restore procedure took so long because they had to make backups first, 5) the restore…

I for one am glad they released a factual account and timeline of what went wrong. I don't see it as an attempt to minimize their actions. They even admit that they have no clear explanation of the original issue, when they could easily have committed to a stronger theory to make themselves look more competent. Overall I'd much rather read this than a massaged PR apology that keeps us in the dark of what actually happened.

Re: Postmortem of the failure of one hosting storage unit on Jan. 8, 2020

#38
post #30

Earlier quoted context omitted.

Fair enough. That probably is bad marketing if nothing else, and maybe something else. what you're saying about the major failure being in support/customer-management, even more than the technical issue, seems potentially reasonable. (I am not a Gandi customer, so it's not personal for me). I still think the fact that there was no data loss, and we're still on the edge of calling it unacceptable incompetence, is wort…

The linked twitter thread explicitly mentions data loss: > Hi Andrea. It is confirmed we have lost data and we are terribly sorry for that. However, please note that what happenend[sp] could happen to any web host. Customers that were forced to migrate to a different webhost had to restore from whatever backups they had, and they lost data for sure. Even if Gandi ultimately recovered everything (and it's not complete…

OK, that makes it even worse then. The postmortem linked above definitely says:

> We managed to restore the data and bring services back online the morning of January 13.

Is that wrong? It's bad to lose data, it's even worse to tell people you didn't lose data in one place when you did, and tell them you did lose data in another.

Re: Postmortem of the failure of one hosting storage unit on Jan. 8, 2020

#39
post #17

They don't say they're sorry, because they're not. Instead they minimize their actions by: 1) stating how few customers customers were affected, 2) how it's not really their fault because it was a hardware error, 3) it's not really their fault because they had already planned to upgrade the server, 4) it's not really their fault the restore procedure took so long because they had to make backups first, 5) the restore…

I for one am glad they released a factual account and timeline of what went wrong. I don't see it as an attempt to minimize their actions. They even admit that they have no clear explanation of the original issue, when they could easily have committed to a stronger theory to make themselves look more competent. Overall I'd much rather read this than a massaged PR apology that keeps us in the dark of what actually hap…

This is the messaged PR “postmortem”. It’s basically a shoulder shrug emoji and takes zero responsibility for the incident.

They also failed to address their abysmal responses on Twitter that essentially belittled and poked fun at the affected users.

E.g. https://news.ycombinator.com/item?id=22002258

Re: Postmortem of the failure of one hosting storage unit on Jan. 8, 2020

#40
post #35

Being a Gandi customer must be terrifying, generally.

They used to have a great reputation in France about 20 to 15 years ago. It went downhill since, the company got sold and they started to sell expensive and slow cloud services, a bit like Amazon with less success.
Post reply on HN