Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

201–210 of 1001 posts

Re: AWS us-east-1 outage

#201

I worked at a company that hired an ex-Amazon engineer to work on some cloud projects. Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault. He later explained that in his…

Perhaps reward structure should be changed to incentivize the post-mortems. There could be several flaws that run underreported otherwise.

We may run into the problem of everything documented and possible deliberate acts but for a service that relies heavily on uptime, that’s a small price to pay for a bulletproof operation.

Re: AWS us-east-1 outage

#203
I'm now getting failures searching for products on Amazon.com itself. This is somewhat surprising, as the narrative always was that Amazon didn't do a great job of dogfooding their own cloud platform.

Re: AWS us-east-1 outage

#204
post #166

I worked at a company that hired an ex-Amazon engineer to work on some cloud projects. Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault. He later explained that in his…

Sometimes, these large companies tack on too much "necessary" incident "remediation" actions with Arbitrary Due Date SLAs that completely wrench any ongoing work. And ongoing, strategically defined ""muh high impact"" projects are what get you promoted, not doing incident remediations. When you get to the level you want, you get to not really give a shit and actually do The Right Thing. However, for all of the engine…

Politicized cloud meh.

Re: AWS us-east-1 outage

#207

I worked at a company that hired an ex-Amazon engineer to work on some cloud projects. Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault. He later explained that in his…

This gets posted every time there's an AWS outage. It mind as well be a copy pasta at this point.

I had that deja vu feeling reading PragmaticPulp's comment, too.

And sure enough, PragmaticPulp did post a similar comment on a thread about Amazon India's alleged hire-to-fire policy 6 months back: https://news.ycombinator.com/item?id=27570411

You and I, we aren't among the 10000, but there are potentially 10000 others who might be: https://xkcd.com/1053/

Re: AWS us-east-1 outage

#208
post #13

I love that every time this happens, 100% of the services on https://status.aws.amazon.com are green.

I wonder if the other parts of Amazon do this. Like their inventory system thinks something is in stock, but people can't find it in the warehouse, do they just simply not send it to you and hope you don't notice? AWS's culture sounds super broken.

My favorite status page, though, is Slack's. You can read an article in the New York Times about how Slack was down for most of a day, and the status page is just like "some percentage of users experienced minor connectivity issues". "Some percentage" is code for "100%" and "minor" is code for "total". Good try.

Re: AWS us-east-1 outage

#209
post #38

Earlier quoted context omitted.

Those five 9s don't come easy. Sometimes you have to prop them up :)

I wonder how often outages really happen. The official page is nonsense, of course, and we only collectively notice when the outage is big enough that lots of us are affected. On AWS, I see about a 3:1 ratio of "bump in the night" outages (quickly resolved, little corroboration) to mega too-big-to-hide outages. Does that mirror others' experiences?

If you count any time AWS is having a problem that impacts our production workloads then I think it's about 5:1. Dealing with "AWS is down" outages are easy because I can just sit back and grab some popcorn, it's the "dammit I know this is AWS's fault" outages that are a PITA because you count yourself lucky to even get a report in your personalized dashboard.

Re: AWS us-east-1 outage

#210
post #90

Earlier quoted context omitted.

Man, some conclusions are being _jumped_ to by this reply.

There is a very long history of US-east-1 being horrible. Just bad. We've told every client we can to get out of there. It's one of the oldest amazon regions, and I think too much old legacy and weird stuff happens there. Use US-west-2.

Isn't us-east-1 where they deploy everything first? And the only region that has 100% of all available services?
Post reply on HN