Earlier quoted context omitted.
That's opposite of my experience at AWS. It's likely that the culture at AWS has changed over the past few years, it's also likely that there's a difference in culture between teams.
Having talked to dozens of Amazon engineers, the only consistent picture I've formed in my head is that the culture varies wildly between teams. The folks on the happiest teams are always aghast at hearing the horror stories.
AWS Cognito is having issues and health dashboards are still green
261–270 of 369 posts
Re: AWS Cognito is having issues and health dashboards are still green
#262Five hours later and nothing has changed. For a company like Amazon this should be unacceptable. Before someone replies and says use a different AZ, that's not possible for everyone. If you use a 3rd party service that is hosted on us-east-1 you can't do anything about it. For example, many Heroku services are broken because of this.
Seriously, I get that something falls over. But to have it be 5 hours to recover, for a service this critical is nuts
Re: AWS Cognito is having issues and health dashboards are still green
#263I'm getting tired of that bullshit. Just admit it.
AWS' conditions for what healthy is and everyone elses is completely different. Kinda makes me wonder what their internals are like.
Re: AWS Cognito is having issues and health dashboards are still green
#264Earlier quoted context omitted.
Not at AWS, retail Amazon, but what I saw was COEs were either normal business process or PIP material depending on which org you worked for. And sometimes just the excuse to get you gone. Where I was about 99.9% of the COEs where just a lesson learned and new process to prevent it. There was one that was basically used as a tool to remove a VERY good engineer, that didn't mesh well with new leadership. A sister org,…
What is PIP?
Re: AWS Cognito is having issues and health dashboards are still green
#265Earlier quoted context omitted.
I think the deeper problem is the interconnectivity between services and their apis. It's too complicated to maintain...
Yeah, the cascading failure of all the other services is a deep architectural issue. Having lots of services that do one thing and one thing well makes a lot of sense. Breaking them out into separate components brings a level of visibility into the system. And it's AWS's whole business model. But it does mean that, fundamentally, service X is available when and only when (WAOW?) services A, B, C, etc. are all availab…
Re: AWS Cognito is having issues and health dashboards are still green
#266Earlier quoted context omitted.
Have worked at AWS before, and I can attest to this. Whenever we had an outage, our director and senior manager would take a call on whether to update the dashboard or not. Having 'red' dashboard catches lot of eyes, so people responsible for making this decision always look at it from political point of view. As a dev oncall, we used to get 20 sev2s per day (an oncall ticket which needs to be handled within 15 mins)…
Wow. If I were in charge, the team running a service should not be the same team who decides whether a given service is healthy. This is pretty damaging info about the unprofessional way AWS actually appears to be run.
Re: AWS Cognito is having issues and health dashboards are still green
#267Earlier quoted context omitted.
At amazon, admitting to a problem will guaranteed lead to having to open a COE, correction of error, which means meetings with executives, inevitable "least effective" rating, development plan, scapegoating, PIP, and firing.
What the hell is a COE? I hate that nobody seems to bother defining their acronyms anymore.
Re: AWS Cognito is having issues and health dashboards are still green
#268Five hours later and nothing has changed. For a company like Amazon this should be unacceptable. Before someone replies and says use a different AZ, that's not possible for everyone. If you use a 3rd party service that is hosted on us-east-1 you can't do anything about it. For example, many Heroku services are broken because of this.
All on the eve of thanksgiving.
Re: AWS Cognito is having issues and health dashboards are still green
#269> It's not posted on SHD as the issue has impacted our ability to post there. Is that not a massive catch-22 for a service dashboard?
Reminds me of a recent outage from IBM Cloud, where the VPN was hosted on IBM Cloud so employees couldn’t log in to fix it, and the email is hosted on IBM Cloud so support teams couldn’t email customers to let them know and even access to their Twitter was behind the non-functional VPN so they couldn’t tweet during the outage either.
Do you have a link for more details?
Re: AWS Cognito is having issues and health dashboards are still green
#270We hired an engineer out of Amazon AWS at a previous company. Whenever one of our cloud services went down, he would go to great lengths to not update our status dashboard. When we finally forced him to update the status page, he would only change it to yellow and write vague updates about how service might be degraded for some customers. He flat out refused to ever admit that the cloud services were down. After some…
Eventually, anyone in that role would get fired. No service has an established 100% availability uptime when measured over its complete existence (welcome to any assertions challenging this, if anyone has any).