Live data from Hacker News

Outage post-mortem

tech.dropbox.com

21–30 of 44 posts

Re: Outage post-mortem

#21
post #6

This is, sadly, not a great post-mortem. They missed an opportunity for goodwill. I don't feel more confident in their level of understanding or ability to remediate the problems that led to it after having read it. I know they have an excellent engineering and operations staff -- this post-mortem doesn't reinforce that, though. A few of the things that jumped out at me after one reading: 1. The apology is the next t…

The thing I always look for in post-mortems is an understanding of the failure of human systems. The technical failures are interesting, but it is the human systems that produced the technical failures. And will keep on producing other failures unless changed.

I hasten to add that I'm not looking for finger-pointing or blame. In retrospectives, I think it's always best to assume that individuals did the best with what they had. [1] But I think it'd be great if Dropbox asked themselves things like "How did we miss this bug?" and "How could have we discovered this recovery issue before it was on the critical path for a public outage?" Questions like that help you solve not just this bug, but all the related latent bugs that you got the same way you got the one that just blew up.

[1] A lesson I learned from Norm Kerth: http://www.retrospectives.com/pages/retroPrimeDirective.html

Re: Outage post-mortem

#23
post #21
post #6

This is, sadly, not a great post-mortem. They missed an opportunity for goodwill. I don't feel more confident in their level of understanding or ability to remediate the problems that led to it after having read it. I know they have an excellent engineering and operations staff -- this post-mortem doesn't reinforce that, though. A few of the things that jumped out at me after one reading: 1. The apology is the next t…

The thing I always look for in post-mortems is an understanding of the failure of human systems. The technical failures are interesting, but it is the human systems that produced the technical failures. And will keep on producing other failures unless changed. I hasten to add that I'm not looking for finger-pointing or blame. In retrospectives, I think it's always best to assume that individuals did the best with wha…

"How did we miss this bug?" and "How could have we discovered this recovery issue before it was on the critical path for a public outage?"

I am pretty sure they would have done that - just that they did not include in the post mortem.

Re: Outage post-mortem

#24
post #14

Earlier quoted context omitted.

I agree with many of your points and appreciate your technical assessment of the actual post-mortem aspect, but your first comment seems particularly nit picky. It's a growing trend that when a company or person fucks up, we expect a big, grandiose, sobbing apology (and when they don't, we blow a gasket - a la Snapchat). Now, I'm not saying that I don't expect companies to be forthright and take ownership of their mi…

I'm admittedly being nit-picky because I feel very strongly about the importance of outage communication. Good communication both during and after an incident can make a tremendous amount of difference in how you are perceived. They decided that it was worth apologizing for near the end of the post. All I'm suggesting is that moving that up near the top and acknowledging up front that they let customers down would ha…

Fair enough. Your comment was more of a spark of a sentiment I've been carrying around for a little while. I can't agree enough that proper outage communication is important.

Re: Outage post-mortem

#25
post #18

Earlier quoted context omitted.

> There's no discussion of the human factors like how the recovery process went, how this issue was missed in testing, or what changes if any they think they should make to their incident response process. Not every company is into that whiney startup blood and tears thing. Those "we( ) worked non-stop for the last 72 hours" often sound a bit desperate. ( ) And by "we", the PR people usually mean the engineers.

That's not at all what I was getting at. It's not about patting yourself on the back or trying to make the team look like heroes. I was more wondering how the mechanics of their incident response processes were managed and whether they planned to make any changes as a result of the review of this incident. Technical remediations are all well and good, but organizational, cultural, and even procedural changes are ofte…

I don't see the need for any of that. What does it really matter they had "incident fatigue"? I don't really care about their internal comms or escalation procedures. If I was a customer, I'd want to know what they are doing to mitigate a similar incident (which they answered), and an apology.

If I wanted a credit, or SLAs weren't met, then I'd talk directly to an account manager.

Re: Outage post-mortem

#26

Earlier quoted context omitted.

Would be nice to have a service who stepped in when shit like this happened. I'd pay good money to have a tiger team appear out of thin air when the shit hit the fan.

I mean customer facing.

Meaning a public relations function? That's what this boils down to, ultimately; a person(s) who know the audience, understands what concerns and questions they have and provides timely answers to them.

I thought their response struck the appropriate level of detail. I don't care to know the inner workings of their processes, but I'd like some indication that they care and that they're working on it. I got that from this.

Re: Outage post-mortem

#27
post #21

Earlier quoted context omitted.

The thing I always look for in post-mortems is an understanding of the failure of human systems. The technical failures are interesting, but it is the human systems that produced the technical failures. And will keep on producing other failures unless changed. I hasten to add that I'm not looking for finger-pointing or blame. In retrospectives, I think it's always best to assume that individuals did the best with wha…

"How did we miss this bug?" and "How could have we discovered this recovery issue before it was on the critical path for a public outage?" I am pretty sure they would have done that - just that they did not include in the post mortem.

The first was answered in the postmortem. The second is something done in time - either it's hard to answer in detail without revealing confidential information, or they are working towards it in the medium term.

Re: Outage post-mortem

#29
post #6

This is, sadly, not a great post-mortem. They missed an opportunity for goodwill. I don't feel more confident in their level of understanding or ability to remediate the problems that led to it after having read it. I know they have an excellent engineering and operations staff -- this post-mortem doesn't reinforce that, though. A few of the things that jumped out at me after one reading: 1. The apology is the next t…

This post was just an incident review for a technology audience. Dropbox posted a separate apology to their users on their main blog: https://blog.dropbox.com/2014/01/back-up-and-running/ . The tone and detail seem totally appropriate since it ran concurrently with the other post.

Re: Outage post-mortem

#30
post #6

This is, sadly, not a great post-mortem. They missed an opportunity for goodwill. I don't feel more confident in their level of understanding or ability to remediate the problems that led to it after having read it. I know they have an excellent engineering and operations staff -- this post-mortem doesn't reinforce that, though. A few of the things that jumped out at me after one reading: 1. The apology is the next t…

This is a silly fetishizing of 'post mortum'

I will think no less of any company with solid technology that experiences a failure and puts an honest effort in communicating a post mortum explanation which is exactly what happened here.

I will, though, lose some respect for people who quibble about perceived faux pas of the explanation because it's losing sight of what is actually important.

Post reply on HN