Live data from Hacker News

AWS Cognito is having issues and health dashboards are still green

status.aws.amazon.com

261–270 of 369 posts

Re: AWS Cognito is having issues and health dashboards are still green

#261
post #216

Earlier quoted context omitted.

That's opposite of my experience at AWS. It's likely that the culture at AWS has changed over the past few years, it's also likely that there's a difference in culture between teams.

Having talked to dozens of Amazon engineers, the only consistent picture I've formed in my head is that the culture varies wildly between teams. The folks on the happiest teams are always aghast at hearing the horror stories.

800,000 person company has to vary from team to team

Re: AWS Cognito is having issues and health dashboards are still green

#262

Five hours later and nothing has changed. For a company like Amazon this should be unacceptable. Before someone replies and says use a different AZ, that's not possible for everyone. If you use a 3rd party service that is hosted on us-east-1 you can't do anything about it. For example, many Heroku services are broken because of this.

Seriously, I get that something falls over. But to have it be 5 hours to recover, for a service this critical is nuts

Pretty critical to other Amazon services too. We use Merchant Web Services to import orders from Amazon. Down since 9:30 AM. At this point we have thousands of orders we are unable to import and process.

Re: AWS Cognito is having issues and health dashboards are still green

#263

I'm getting tired of that bullshit. Just admit it.

AWS' conditions for what healthy is and everyone elses is completely different. Kinda makes me wonder what their internals are like.

Probably similarly to the "Downfall" Hitler parody, or the "he's delusional, take him to the infirmary" from Chernobyl, take your pick.

Re: AWS Cognito is having issues and health dashboards are still green

#264

Earlier quoted context omitted.

Not at AWS, retail Amazon, but what I saw was COEs were either normal business process or PIP material depending on which org you worked for. And sometimes just the excuse to get you gone. Where I was about 99.9% of the COEs where just a lesson learned and new process to prevent it. There was one that was basically used as a tool to remove a VERY good engineer, that didn't mesh well with new leadership. A sister org,…

What is PIP?

Performance Improvement Plan. In theory, it sounds like a plan to fix your supposedly inadequate performance. In practice, like 99% of the time it means somebody decided to fire you for some reason before you have even seen the first one, and they're just creating documentation for why they fired you to head off HR requirements and any future complaints. They'll run you through a few rounds of supposedly evaluating your improvements as inadequate and eventually fire you, unless you quit first.

Re: AWS Cognito is having issues and health dashboards are still green

#265
post #228

Earlier quoted context omitted.

I think the deeper problem is the interconnectivity between services and their apis. It's too complicated to maintain...

Yeah, the cascading failure of all the other services is a deep architectural issue. Having lots of services that do one thing and one thing well makes a lot of sense. Breaking them out into separate components brings a level of visibility into the system. And it's AWS's whole business model. But it does mean that, fundamentally, service X is available when and only when (WAOW?) services A, B, C, etc. are all availab…

Interested to know what the alternative might be and why it would mean better uptime

Re: AWS Cognito is having issues and health dashboards are still green

#266
post #195

Earlier quoted context omitted.

Have worked at AWS before, and I can attest to this. Whenever we had an outage, our director and senior manager would take a call on whether to update the dashboard or not. Having 'red' dashboard catches lot of eyes, so people responsible for making this decision always look at it from political point of view. As a dev oncall, we used to get 20 sev2s per day (an oncall ticket which needs to be handled within 15 mins)…

Wow. If I were in charge, the team running a service should not be the same team who decides whether a given service is healthy. This is pretty damaging info about the unprofessional way AWS actually appears to be run.

It's funny that you point to that as the problem. The problem is more AWS' toxic engineering culture that has engineers fearing for their jobs in a way that guides their decision making. It's bad company culture, end of story.

Re: AWS Cognito is having issues and health dashboards are still green

#267

Earlier quoted context omitted.

At amazon, admitting to a problem will guaranteed lead to having to open a COE, correction of error, which means meetings with executives, inevitable "least effective" rating, development plan, scapegoating, PIP, and firing.

What the hell is a COE? I hate that nobody seems to bother defining their acronyms anymore.

https://wa.aws.amazon.com/wat.concept.coe.en.html

Re: AWS Cognito is having issues and health dashboards are still green

#268

Five hours later and nothing has changed. For a company like Amazon this should be unacceptable. Before someone replies and says use a different AZ, that's not possible for everyone. If you use a 3rd party service that is hosted on us-east-1 you can't do anything about it. For example, many Heroku services are broken because of this.

I can imagine that there are literally 100s of engineers involved in trying to fix this ASAP, since this is not only bringing down the systems of external customers, but also critical internal systems, plus the bad PR.

All on the eve of thanksgiving.

Re: AWS Cognito is having issues and health dashboards are still green

#269
post #8

> It's not posted on SHD as the issue has impacted our ability to post there. Is that not a massive catch-22 for a service dashboard?

Reminds me of a recent outage from IBM Cloud, where the VPN was hosted on IBM Cloud so employees couldn’t log in to fix it, and the email is hosted on IBM Cloud so support teams couldn’t email customers to let them know and even access to their Twitter was behind the non-functional VPN so they couldn’t tweet during the outage either.

This sounds like a true nightmare.

Do you have a link for more details?

Re: AWS Cognito is having issues and health dashboards are still green

#270

We hired an engineer out of Amazon AWS at a previous company. Whenever one of our cloud services went down, he would go to great lengths to not update our status dashboard. When we finally forced him to update the status page, he would only change it to yellow and write vague updates about how service might be degraded for some customers. He flat out refused to ever admit that the cloud services were down. After some…

Eventually, anyone in that role would get fired. No service has an established 100% availability uptime when measured over its complete existence (welcome to any assertions challenging this, if anyone has any).

Bitcoin?
Post reply on HN