Live data from Hacker News

AWS Cognito is having issues and health dashboards are still green

status.aws.amazon.com

241–250 of 369 posts

Re: AWS Cognito is having issues and health dashboards are still green

#241

Now is probably a good time to plug some of the open source alternatives to vendor locked in identity solutions: - https://github.com/ory - https://github.com/dexidp/dex - https://github.com/authelia/authelia - https://github.com/keycloak/keycloak - https://www.gluu.org/ - https://github.com/accounts-js/accounts

im surprised companies still want to build their own identity system or pay companies (ping, auth0) to host it for them

ory looks like a really good project

Re: AWS Cognito is having issues and health dashboards are still green

#242

Earlier quoted context omitted.

Not at AWS, retail Amazon, but what I saw was COEs were either normal business process or PIP material depending on which org you worked for. And sometimes just the excuse to get you gone. Where I was about 99.9% of the COEs where just a lesson learned and new process to prevent it. There was one that was basically used as a tool to remove a VERY good engineer, that didn't mesh well with new leadership. A sister org,…

What is PIP?

Performance Improvement Plan, they are not unique to Amazon, most places have them though the process may differ. Not to be too cynical but ultimately they’re a way to document that you’re not meeting expectations - before being fired. Should there be any sort of employment claim later its a mechanism by which an employer can show documentation that any issues related to your being let go were performance related and not some sort of protected status or prejudice.

Outside of someone protected by a labor union, I’ve very rarely seen anyone recover from a PIP and not be eventually let go. Most commonly employees see them as a 30 or 60 day window to proactively find a new job before they’re terminated.

Re: AWS Cognito is having issues and health dashboards are still green

#243

Earlier quoted context omitted.

What is PIP?

Amazon fires between 5-15% of engineers per year. PIP is to get you to quit. Amazon hires a TON of entry level SDE 1 engineers to sacrifice at the altar of Bezos so more shitty employees get to stay. Lifespan of a SDE 1 whipping boy/girl at Amazon, as a result is 3-6 months.

Only the strong survive:p

Re: AWS Cognito is having issues and health dashboards are still green

#244

We hired an engineer out of Amazon AWS at a previous company. Whenever one of our cloud services went down, he would go to great lengths to not update our status dashboard. When we finally forced him to update the status page, he would only change it to yellow and write vague updates about how service might be degraded for some customers. He flat out refused to ever admit that the cloud services were down. After some…

I have heard stories like these before but it wasn’t clear to me that this is apparently a broader issue at AWS (reading the other comments). While I think that very short outages in line with SLAs must mot necessarily go public or have a post mortem, it is astonishing to see that some teams/managers go through lengths to hide this at the „primus“ of hyperscalers.

I always wonder how many more products AWS pushes out the door versus cleaning up and improving what the have already. Cognito itself is such a half-baked mess...

But back to topic, when should we update status pages? On every incident? Or when SLAs are violated?

Re: AWS Cognito is having issues and health dashboards are still green

#245
post #221

Earlier quoted context omitted.

No idea what happens on AWS as I don't work there, but I have another perspective on this. There are perverse incentives to NOT update your status dashboard. Once I was asked by management to _take our status dashboard down_ . That sounded backwards, so I dug a bit more. Turns out our competitor was using our status dashboard as ammo against us in their sales pitch. Their claim was that we had too many issues and wer…

Yeah, had the same experience at a previous company. It's very frustrating that your transparency gets used against you by unscrupulous competitors.

How is it unscrupulous?

This sort of shit happens all the time at all levels. Companies use each other’s public specs in their competition all the time.

Or capitalizing on features like headphone jacks etc. in their ads before proceeding to remove them from their own products anyway (Samsung and Google) and so on.

Re: AWS Cognito is having issues and health dashboards are still green

#246

Five hours later and nothing has changed. For a company like Amazon this should be unacceptable. Before someone replies and says use a different AZ, that's not possible for everyone. If you use a 3rd party service that is hosted on us-east-1 you can't do anything about it. For example, many Heroku services are broken because of this.

Seriously, I get that something falls over. But to have it be 5 hours to recover, for a service this critical is nuts

Re: AWS Cognito is having issues and health dashboards are still green

#247
post #195

We hired an engineer out of Amazon AWS at a previous company. Whenever one of our cloud services went down, he would go to great lengths to not update our status dashboard. When we finally forced him to update the status page, he would only change it to yellow and write vague updates about how service might be degraded for some customers. He flat out refused to ever admit that the cloud services were down. After some…

Have worked at AWS before, and I can attest to this. Whenever we had an outage, our director and senior manager would take a call on whether to update the dashboard or not. Having 'red' dashboard catches lot of eyes, so people responsible for making this decision always look at it from political point of view. As a dev oncall, we used to get 20 sev2s per day (an oncall ticket which needs to be handled within 15 mins)…

Wow. If I were in charge, the team running a service should not be the same team who decides whether a given service is healthy. This is pretty damaging info about the unprofessional way AWS actually appears to be run.

Re: AWS Cognito is having issues and health dashboards are still green

#248

Five hours later and nothing has changed. For a company like Amazon this should be unacceptable. Before someone replies and says use a different AZ, that's not possible for everyone. If you use a 3rd party service that is hosted on us-east-1 you can't do anything about it. For example, many Heroku services are broken because of this.

Seriously, I get that something falls over. But to have it be 5 hours to recover, for a service this critical is nuts

More like ten hours at this point for Kinesis

Re: AWS Cognito is having issues and health dashboards are still green

#249

We hired an engineer out of Amazon AWS at a previous company. Whenever one of our cloud services went down, he would go to great lengths to not update our status dashboard. When we finally forced him to update the status page, he would only change it to yellow and write vague updates about how service might be degraded for some customers. He flat out refused to ever admit that the cloud services were down. After some…

Eventually, anyone in that role would get fired. No service has an established 100% availability uptime when measured over its complete existence (welcome to any assertions challenging this, if anyone has any).

Re: AWS Cognito is having issues and health dashboards are still green

#250

Earlier quoted context omitted.

Amazon fires between 5-15% of engineers per year. PIP is to get you to quit. Amazon hires a TON of entry level SDE 1 engineers to sacrifice at the altar of Bezos so more shitty employees get to stay. Lifespan of a SDE 1 whipping boy/girl at Amazon, as a result is 3-6 months.

Only the strong survive:p

Shit floats :p
Post reply on HN