Live data from Hacker News

Post-mortem of Christmas Eve Amazon ELB outage

aws.amazon.com

1–10 of 40 posts

Re: Post-mortem of Christmas Eve Amazon ELB outage

#4
post #2

"The data was deleted by a maintenance process that was inadvertently run against the production ELB state data. This process was run by one of a very small number of developers who have access to this production environment."

And the developer ought to be fired. He cost me a considerable amount of money since my heroku site with SSL endpoint was down over 12 hours during one if our biggest days of the year. Of course, heroku still (over)bills me those hours for Dynos and Workers.

I course, I am small potatoes compared to Netflix being down (while Amazon instant video wasn't.)

Amazon is populated by jackasses apparently and Heroku isn't much better. Why the f*(k doesn't Heroku have some kind of fall-over or something to protect clients from the inevitable AWS-East failures?

Buh, bye Heroku and AWS. Hello Bluebox. I run other apps on Bluebox and have never had any downtime except for some mild nonsense due to my own damned ignorance. Luckily the Blueboxers saved me from myself. AWS needs to get their shit together.

Re: Post-mortem of Christmas Eve Amazon ELB outage

#6
post #2

"The data was deleted by a maintenance process that was inadvertently run against the production ELB state data. This process was run by one of a very small number of developers who have access to this production environment."

And the developer ought to be fired. He cost me a considerable amount of money since my heroku site with SSL endpoint was down over 12 hours during one if our biggest days of the year. Of course, heroku still (over)bills me those hours for Dynos and Workers. I course, I am small potatoes compared to Netflix being down (while Amazon instant video wasn't.) Amazon is populated by jackasses apparently and Heroku isn't mu…

In my experience, most really great developers have made similar mistakes at some point in their pasts. Often these mistakes are what catalyze growth and create the scar tissue that helps them (and organizations) to not repeat the same classes of mistakes. You don't fire people when they make mistakes like these, you ask them what they will do to make sure that nobody in the organization is likely to make a similar one. Seems like that is generally Amazon's reaction too.

That said, I understand it's pretty damn frustrating having downtime during peak periods :)

Re: Post-mortem of Christmas Eve Amazon ELB outage

#7
post #5

It's amazing that the team worked through what is likely the least fun night of the year to be working to fix this issue.

I wonder who gets stuck working those shifts -- would it just be following the normal schedule, in the name of fairness? Do more senior people get to claim the day of vacation? Do parents get priority over single people?

Re: Post-mortem of Christmas Eve Amazon ELB outage

#10
post #2

"The data was deleted by a maintenance process that was inadvertently run against the production ELB state data. This process was run by one of a very small number of developers who have access to this production environment."

And the developer ought to be fired. He cost me a considerable amount of money since my heroku site with SSL endpoint was down over 12 hours during one if our biggest days of the year. Of course, heroku still (over)bills me those hours for Dynos and Workers. I course, I am small potatoes compared to Netflix being down (while Amazon instant video wasn't.) Amazon is populated by jackasses apparently and Heroku isn't mu…

You weren't down because an Amazon engineer fat-fingered a maintenance command, you were down because you chose to put all of your eggs in one basket. The basket you chose failed, and you had no redundancy basket to fail over to. Choosing a new basket to put all your eggs in isn't going to fix anything just because the new basket hasn't failed on you yet.

Make smarter decisions. If being up on December 24 was really that important to you, you'd have had backup hosting in place with the ability to quickly fail over to it. You'll start being a better engineer when you quit blaming some poor bastard's bad luck for your failures and learn the real lesson from your downtime: you fucked up by not having high-availability hosting to meet your claimed high-availability needs.

The worst thing Amazon could do is fire this guy. They've just paid a lot of money to have him learn from his mistake. They'd be fools to throw that hard-earned experience away.

Post reply on HN