Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

211–220 of 1001 posts

Re: AWS us-east-1 outage

#211

I worked at a company that hired an ex-Amazon engineer to work on some cloud projects. Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault. He later explained that in his…

That’s idiotic, the service is down regardless. If you foster that kind of culture, why have a status page at all?

It make AWS engineers look stupid, because it looks like they are not monitoring their services.

Re: AWS us-east-1 outage

#212
A former colleague told me years ago that us-east-1 is basically the guinea pig where changes get tested before being rolled out to the other regions, and as a result is less stable than the others. Does anyone know if there's any truth to this?

Re: AWS us-east-1 outage

#213
Not sure it helps, but got this update from someone inside AWS a few moments ago.

"We have identified the root cause of the issues in the US-EAST-1 Region, which is a network issue with some network devices in that Region which is affecting multiple services, including the console but also services like S3. We are actively working towards recovery."

Re: AWS us-east-1 outage

#214
post #113

Friends tell friends to pick us-east-2. Virginia is for lovers, Ohio is for availability.

If you're not multi-cloud in 2021 and are expecting 5-9's, I feel bad for you.

If you're having SLA problems I feel bad for you son

I got two 9 problems cuz of us-east-1

Re: AWS us-east-1 outage

#215
I am able to access the console for us-west-2 by going to a region specific URL: https://us-west-2.console.aws.amazon.com/

It does tack me to a non-region specific login page which is up, and then redirects back to us-west-2 which works.

If I go to my bookmarked login page it breaks because, presumably, it hits something that is broken in us-east-1.

Re: AWS us-east-1 outage

#217

I worked at a company that hired an ex-Amazon engineer to work on some cloud projects. Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault. He later explained that in his…

That sounds like the exact opposite of human-factors engineering. No one likes taking blame. But when things go sideways, people are extra spicy and defensive, which makes them clam up and often withhold useful information, which can extend the outage.

No-blame analysis is a much better pattern. Everyone wins. It's about building the system that builds the system. Stuff broke; fix the stuff that broke, then fix the things that let stuff break.

Re: AWS us-east-1 outage

#218

I worked at a company that hired an ex-Amazon engineer to work on some cloud projects. Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault. He later explained that in his…

Damn, he had serious PTSD lol

Re: AWS us-east-1 outage

#220
post #201

I worked at a company that hired an ex-Amazon engineer to work on some cloud projects. Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault. He later explained that in his…

Perhaps reward structure should be changed to incentivize the post-mortems. There could be several flaws that run underreported otherwise. We may run into the problem of everything documented and possible deliberate acts but for a service that relies heavily on uptime, that’s a small price to pay for a bulletproof operation.

Then we would drown in a sea of meetings and 'lessons learned' emails. There is a reason for post-mortems, but there has to be balance.
Post reply on HN