Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

341–350 of 1001 posts

Re: AWS us-east-1 outage

#341
post #274

It's funny that the first place I go to learn about the outage is Hacker News and not https://status.aws.amazon.com/ (it's still reports everything to be "operating normally"...)

Now 57 minutes later and it still reports everything as operating normally.

It shows errors now.

Re: AWS us-east-1 outage

#342
So, we're getting failures (for customers) trying to use amazon pay from our site. AFAIK there is no "status page" for Amazon Pay, but the rest of Amazon's services seem to be a giant Rube Goldberg machine so it's hard to imagine this isn't too.

Re: AWS us-east-1 outage

#343
post #119

Earlier quoted context omitted.

Definitely not just the console. We had hundreds of thousands of websocket connections to us-east-1 drop at 15:40, and new websocket connections to that region are still failing. (Luckily not a huge impact on our service cause we run in 6 other regions, but still).

Side question: How happy are you with API Gateway's WebSocket service?

No idea, we don't use it. These were websocket connections to processes on ec2, via NLB and cloudfront. Not sure exactly what part of that chain was broken yet.

Re: AWS us-east-1 outage

#344

Earlier quoted context omitted.

Seems "the cloud" had a major outage less than a month ago, my laptop has a higher uptime. $ 16:04 up 46 days, 7:02, 9 users, load averages: 3.68 3.56 3.18 US East 1 was down just over a year ago https://www.theregister.com/2020/11/25/aws_down/ Meanwhile I moved one of my two internal DNS servers to a second site on 11 Nov 2020, and it's been up since then. One of my monitoring machines has been filling, rotating and…

I think the point of the could isn't increased uptime - the point is that when it's down, bring it back up is someone else's problem . (Also, OpEx vs CapEx financial shenanigans...) All the same, I don't disagree with your point.

> the point is that when it's down, bring it back up is someone else's problem.

When it's down, it's my problem, and I can't do anything about it other than explain why I have no idea the system is broken and can't do anything about it.

"Why is my dohicky down? When will it be back?"

"Because it's raining, no idea"

May be accurate, it's also of no use.

But yes, Opex vs Capex, of course that's why you can lease your servers. It's far easier to spend company money with another $500 a month on AWS than spend $500 a year for a new machine.

Re: AWS us-east-1 outage

#346
post #217

Earlier quoted context omitted.

That sounds like the exact opposite of human-factors engineering. No one likes taking blame. But when things go sideways, people are extra spicy and defensive, which makes them clam up and often withhold useful information, which can extend the outage. No-blame analysis is a much better pattern. Everyone wins. It's about building the system that builds the system. Stuff broke; fix the stuff that broke, then fix the t…

I worked at Walmart Technology. I bravely wrote post mortem documents owning the fault of my team (100+ people), owning both technically and also culturally as their leader. I put together a plan to fix it and executed it. Thought that was the right thing to do. This happend two times in my 10 year career there. Both times I was called out as a failure in my performance eval. Second time, I resigned and told them to…

Props to you and Walmart will never realize their loss. Unfortunately. But one day there will be headline (or even a couple of them) and you will know that if you had been there it might not have happened and that in the end it is Walmarts' customers that will pay the price for that, not their shareholders.

Re: AWS us-east-1 outage

#347
post #122

Earlier quoted context omitted.

Sure, but... that just raises more questions :) Taken literally what you are saying is the service could be down and an executive could override that, preventing them for paying customers for a service outage, even if the service did have an outage and the customer could prove it (screenshots, metrics from other cloud providers, many different folks see it). I'm sure there is some subtlety to this, but it does mean t…

Large corps with influence get what they want regardless. Status page goes red and the small corps start thinking they can get what they want too.

> Status page goes red and the small corps start thinking they can get what they want too.

I think you mean "start thinking they can get what they pay for"

Re: AWS us-east-1 outage

#348
post #212

A former colleague told me years ago that us-east-1 is basically the guinea pig where changes get tested before being rolled out to the other regions, and as a result is less stable than the others. Does anyone know if there's any truth to this?

I can't see why they'd use the most common/popular region as a guinea pig.

Re: AWS us-east-1 outage

#349
post #308

I worked at a company that hired an ex-Amazon engineer to work on some cloud projects. Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault. He later explained that in his…

Haha... This bring back memories. It really depends on the org. I've had push backs on my postmortems before because of phrasing that could be constituted as laying some of the blame on some person/team when it's supposed to be blameless. And for a long time, it was fairly blameless. You would still be punished with the extra work of writing high quality postmortems, but I have seen people accidentally bring down cri…

We're in a situation where the balls of mud made people afraid to touch some things in the system. As experiences and processes have improved we've started to crack back into those things and guess what, when you are being groomed to own a process you're going to fuck it up from time to time. Objectively, we're still breaking production less often per year than other teams, but we are breaking it, and that's novel behavior, so we have to keep reminding people why.

The moment that affects promotions negatively, or your coworkers throw you under the bus, you should 1) be assertive and 2) proof-read your resume as a precursor to job hunting.

Re: AWS us-east-1 outage

#350

I worked at a company that hired an ex-Amazon engineer to work on some cloud projects. Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault. He later explained that in his…

That’s idiotic, the service is down regardless. If you foster that kind of culture, why have a status page at all? It make AWS engineers look stupid, because it looks like they are not monitoring their services.

> It make AWS engineers look stupid, because it looks like they are not monitoring their services.

Management.

Post reply on HN