Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

271–280 of 1001 posts

Re: AWS us-east-1 outage

#271
post #245

Looks like they've acknowledged it on the status page now. https://status.aws.amazon.com/ > 8:22 AM PST We are investigating increased error rates for the AWS Management Console. > 8:26 AM PST We are experiencing API and console issues in the US-EAST-1 Region. We have identified root cause and we are actively working towards recovery. This issue is affecting the global console landing page, which is also hosted in US…

They are still lying about it, the issues are not only affecting the console but also AWS operations such as S3 puts. S3 still shows green.

Re: AWS us-east-1 outage

#272

It's funny that the first place I go to learn about the outage is Hacker News and not https://status.aws.amazon.com/ (it's still reports everything to be "operating normally"...)

Community reporting > internal operations

Re: AWS us-east-1 outage

#273
post #245

Looks like they've acknowledged it on the status page now. https://status.aws.amazon.com/ > 8:22 AM PST We are investigating increased error rates for the AWS Management Console. > 8:26 AM PST We are experiencing API and console issues in the US-EAST-1 Region. We have identified root cause and we are actively working towards recovery. This issue is affecting the global console landing page, which is also hosted in US…

https://status.aws.amazon.com/ still shows all green for me

It's acting odd for me. Shows all green in Firefox, but shows the error in Chrome even after some refreshes. Not sure what's caching where to cause that.

Re: AWS us-east-1 outage

#274

It's funny that the first place I go to learn about the outage is Hacker News and not https://status.aws.amazon.com/ (it's still reports everything to be "operating normally"...)

Now 57 minutes later and it still reports everything as operating normally.

Re: AWS us-east-1 outage

#275
post #217

Earlier quoted context omitted.

That sounds like the exact opposite of human-factors engineering. No one likes taking blame. But when things go sideways, people are extra spicy and defensive, which makes them clam up and often withhold useful information, which can extend the outage. No-blame analysis is a much better pattern. Everyone wins. It's about building the system that builds the system. Stuff broke; fix the stuff that broke, then fix the t…

Or just take responsibility. People will respect you for doing that and you will demonstrate leadership.

Cynical/realist take: Take responsibility and then hope your bosses already love you, you can immediately both come with a way to prevent it from happening again, and convince them to give you the resources to implement it. Otherwise your responsibility is, unfortunately, just blood in the water for someone else to do all of that to protect the company against you and springboard their reputation on the descent of yours. There were already senior people scheming to take over your department from your bosses, now they have an excuse.

Re: AWS us-east-1 outage

#276
post #243

Earlier quoted context omitted.

Or just take responsibility. People will respect you for doing that and you will demonstrate leadership.

And the guy who doesn't take responsibility gets promoted. Employees are not responsible for failures of management to set a good culture.

Not in healthy organizations, they don't.

Re: AWS us-east-1 outage

#277
post #217

I worked at a company that hired an ex-Amazon engineer to work on some cloud projects. Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault. He later explained that in his…

That sounds like the exact opposite of human-factors engineering. No one likes taking blame. But when things go sideways, people are extra spicy and defensive, which makes them clam up and often withhold useful information, which can extend the outage. No-blame analysis is a much better pattern. Everyone wins. It's about building the system that builds the system. Stuff broke; fix the stuff that broke, then fix the t…

I don't think engineers can believe in no-blame analysis if they know it'll harm career growth. I can't unilaterally promote John Doe, I have to convince other leaders that John would do well the next level up. And in those discussions, they could bring up "but John has caused 3 incidents this year", and honestly, maybe they'd be right.

Re: AWS us-east-1 outage

#278

Earlier quoted context omitted.

Yeah, strange, my self-hosted server isn't affected either.

Seems "the cloud" had a major outage less than a month ago, my laptop has a higher uptime. $ 16:04 up 46 days, 7:02, 9 users, load averages: 3.68 3.56 3.18 US East 1 was down just over a year ago https://www.theregister.com/2020/11/25/aws_down/ Meanwhile I moved one of my two internal DNS servers to a second site on 11 Nov 2020, and it's been up since then. One of my monitoring machines has been filling, rotating and…

I think the point of the could isn't increased uptime - the point is that when it's down, bring it back up is someone else's problem.

(Also, OpEx vs CapEx financial shenanigans...)

All the same, I don't disagree with your point.

Post reply on HN