Live data from Hacker News

Outage post-mortem

tech.dropbox.com

11–20 of 44 posts

Re: Outage post-mortem

#11
post #8
post #6

This is, sadly, not a great post-mortem. They missed an opportunity for goodwill. I don't feel more confident in their level of understanding or ability to remediate the problems that led to it after having read it. I know they have an excellent engineering and operations staff -- this post-mortem doesn't reinforce that, though. A few of the things that jumped out at me after one reading: 1. The apology is the next t…

As I've said before, one blog post does not represent that team when it's written by someone tasked with the job of communicating with a wide variety of customers. My mom could give two hoots about details. She wants to know why her 'spinny drobox thing' keep spinning and should she upgrade or something. I deliver that news to her. This blog post delivers it to people who don't understand as well as most of us but be…

That's just it, though. This is the public face of the team that responded to that outage. It absolutely represents them. Now, whether it's a fair depiction or not is definitely a valid question.

Having written more than my fair share of these, I definitely understand the difficulty involved in choosing your audience and writing to them. That's a big part of the problem here: The audience is not clear. It bounces between technical detail like MySQL recovery process, but it doesn't go deep enough to be satisfying for a really technical audience while being too detailed for a non-technical one.

I have nothing but admiration for their team and the service they've built, but this post-mortem misses the mark.

Re: Outage post-mortem

#12
> The service was back up and running about three hours later, with core service fully restored by 4:40 PM PT on Sunday.

I understand these things happen but I didn't have anything working at all until Sunday EST. I'm just happy it's back.

Having said that, what do you use to backup your Dropbox? I recently signed up for Bitcasa but my Dropbox folder had not been fully uploaded by the time Dropbox stopped worked.

Re: Outage post-mortem

#13
post #4

I found it hard to work out where to get the most up to date information on the outage. I checked the blog, but the last message was their New Year message. In the app and on the main website (mobile version) I couldn't see anything... Glad I read HN otherwise I don't know how I would have come across this information. =)

Would be nice to have a service who stepped in when shit like this happened. I'd pay good money to have a tiger team appear out of thin air when the shit hit the fan.

Like a SWAT team you mean

Re: Outage post-mortem

#14
post #6

This is, sadly, not a great post-mortem. They missed an opportunity for goodwill. I don't feel more confident in their level of understanding or ability to remediate the problems that led to it after having read it. I know they have an excellent engineering and operations staff -- this post-mortem doesn't reinforce that, though. A few of the things that jumped out at me after one reading: 1. The apology is the next t…

I agree with many of your points and appreciate your technical assessment of the actual post-mortem aspect, but your first comment seems particularly nit picky. It's a growing trend that when a company or person fucks up, we expect a big, grandiose, sobbing apology (and when they don't, we blow a gasket - a la Snapchat).

Now, I'm not saying that I don't expect companies to be forthright and take ownership of their mistakes, as well as apologize for them, but I can't help but feeling that expecting Dropbox and others to get on their knees and kiss their users' toes when something happens is a little melodramatic. On the one hand, yes, they made a mistake - on the other, we all know that technology is flawed, and these things happen, albeit rarely.

TL;DR: Let's not make a drama out of it.

Re: Outage post-mortem

#15
post #14
post #6

This is, sadly, not a great post-mortem. They missed an opportunity for goodwill. I don't feel more confident in their level of understanding or ability to remediate the problems that led to it after having read it. I know they have an excellent engineering and operations staff -- this post-mortem doesn't reinforce that, though. A few of the things that jumped out at me after one reading: 1. The apology is the next t…

I agree with many of your points and appreciate your technical assessment of the actual post-mortem aspect, but your first comment seems particularly nit picky. It's a growing trend that when a company or person fucks up, we expect a big, grandiose, sobbing apology (and when they don't, we blow a gasket - a la Snapchat). Now, I'm not saying that I don't expect companies to be forthright and take ownership of their mi…

I'm admittedly being nit-picky because I feel very strongly about the importance of outage communication. Good communication both during and after an incident can make a tremendous amount of difference in how you are perceived.

They decided that it was worth apologizing for near the end of the post. All I'm suggesting is that moving that up near the top and acknowledging up front that they let customers down would have improved the outcome.

They don't need to be over the top about it, just don't bury it at the end of the post.

Re: Outage post-mortem

#16
post #8

Earlier quoted context omitted.

As I've said before, one blog post does not represent that team when it's written by someone tasked with the job of communicating with a wide variety of customers. My mom could give two hoots about details. She wants to know why her 'spinny drobox thing' keep spinning and should she upgrade or something. I deliver that news to her. This blog post delivers it to people who don't understand as well as most of us but be…

That's just it, though. This is the public face of the team that responded to that outage. It absolutely represents them. Now, whether it's a fair depiction or not is definitely a valid question. Having written more than my fair share of these, I definitely understand the difficulty involved in choosing your audience and writing to them. That's a big part of the problem here: The audience is not clear. It bounces bet…

> the audience is not clear.

Bingo. We need nerd updates.

BTW, we deserve this because enough of use use Dropbox for quite important things coding-wise.

Re: Outage post-mortem

#17
post #4

I found it hard to work out where to get the most up to date information on the outage. I checked the blog, but the last message was their New Year message. In the app and on the main website (mobile version) I couldn't see anything... Glad I read HN otherwise I don't know how I would have come across this information. =)

Would be nice to have a service who stepped in when shit like this happened. I'd pay good money to have a tiger team appear out of thin air when the shit hit the fan.

I mean customer facing.

Re: Outage post-mortem

#18
post #6

This is, sadly, not a great post-mortem. They missed an opportunity for goodwill. I don't feel more confident in their level of understanding or ability to remediate the problems that led to it after having read it. I know they have an excellent engineering and operations staff -- this post-mortem doesn't reinforce that, though. A few of the things that jumped out at me after one reading: 1. The apology is the next t…

> There's no discussion of the human factors like how the recovery process went, how this issue was missed in testing, or what changes if any they think they should make to their incident response process.

Not every company is into that whiney startup blood and tears thing. Those "we() worked non-stop for the last 72 hours" often sound a bit desperate.

() And by "we", the PR people usually mean the engineers.

Re: Outage post-mortem

#19
post #18
post #6

This is, sadly, not a great post-mortem. They missed an opportunity for goodwill. I don't feel more confident in their level of understanding or ability to remediate the problems that led to it after having read it. I know they have an excellent engineering and operations staff -- this post-mortem doesn't reinforce that, though. A few of the things that jumped out at me after one reading: 1. The apology is the next t…

> There's no discussion of the human factors like how the recovery process went, how this issue was missed in testing, or what changes if any they think they should make to their incident response process. Not every company is into that whiney startup blood and tears thing. Those "we( ) worked non-stop for the last 72 hours" often sound a bit desperate. ( ) And by "we", the PR people usually mean the engineers.

That's not at all what I was getting at. It's not about patting yourself on the back or trying to make the team look like heroes.

I was more wondering how the mechanics of their incident response processes were managed and whether they planned to make any changes as a result of the review of this incident. Technical remediations are all well and good, but organizational, cultural, and even procedural changes are often even more impactful after events like this.

For example:

Were they happy with the pace of communication during the outage? Do they think customers were updated frequently enough, too frequently, etc. Any changes planned?

How did they handle incident fatigue? Did they have to go to shifts to manage the recovery? Did they already have this planned or was it done on the fly? Do they plan to build any procedures to handle similar long-running events in the future?

Re: Outage post-mortem

#20
agree , give the story . what was the command in the script that failed. to error is human to blog honestly about it is a story I want to read. there is no room for fear in good content , show the true story.
Post reply on HN