Live data from Hacker News

AWS Cognito is having issues and health dashboards are still green

status.aws.amazon.com

201–210 of 369 posts

Re: AWS Cognito is having issues and health dashboards are still green

#201

Now is probably a good time to plug some of the open source alternatives to vendor locked in identity solutions: - https://github.com/ory - https://github.com/dexidp/dex - https://github.com/authelia/authelia - https://github.com/keycloak/keycloak - https://www.gluu.org/ - https://github.com/accounts-js/accounts

Anyone have thoughts on their experience with keycloak?

Re: AWS Cognito is having issues and health dashboards are still green

#203

We hired an engineer out of Amazon AWS at a previous company. Whenever one of our cloud services went down, he would go to great lengths to not update our status dashboard. When we finally forced him to update the status page, he would only change it to yellow and write vague updates about how service might be degraded for some customers. He flat out refused to ever admit that the cloud services were down. After some…

This sounds like a managerial incentives problem.

If a person or company’s compensation depends on not fessing up to problems, they won’t fess up to them.

Re: AWS Cognito is having issues and health dashboards are still green

#204
post #73

Earlier quoted context omitted.

That day was a nightmare for a lot of people - it wasn't just S3 that went down, it was like all of US-EAST. Luckily my company decided against multi-az for the cost savings so I spent all day firefighting.

Multi-AZ doesn't help when a whole region is down, unless you're referring to multi-region AZs (e.g us-east-1a and us-west-1a)

I have to think they're talking about the latter.

Re: AWS Cognito is having issues and health dashboards are still green

#205
post #25

Can anyone explain why status pages are so difficult. Theres even statups like status.io dedicated to this one thing. It really does seem that anytime there is an outage more often than not the status page is showing all green traffic lights. Making it redundant as a tool to corroborate whats happening. How did AWS status page compare with status.io/aws?

This is why my instinct is to check Twitter feeds of the related service first. So far in several years of experience it has been more informative and helpful than a status has ever been. It's a sad state.

One never thought we'd see the day... Twitter, that storied home of the whales of fail, is the reliable service.

Re: AWS Cognito is having issues and health dashboards are still green

#206
post #25

Can anyone explain why status pages are so difficult. Theres even statups like status.io dedicated to this one thing. It really does seem that anytime there is an outage more often than not the status page is showing all green traffic lights. Making it redundant as a tool to corroborate whats happening. How did AWS status page compare with status.io/aws?

It's not that easy to quantify how down a service is at the scale of AWS, for example Cognito has issues, does it means every services that rely on that have issues, what is the impact etc...

Re: AWS Cognito is having issues and health dashboards are still green

#207

Earlier quoted context omitted.

I've caused and authored many COEs at Amazon, and additionally have been involved in maybe fifty for neighboring teams. I can't recall a time it had a negative career impact for anyone, much less any of the consequences you list.

Same. If anything, a well written COE has had positive career impact.

Hah. I don't work for AWS (anymore) but a COE I wrote was literally on my promo doc

Re: AWS Cognito is having issues and health dashboards are still green

#208

We hired an engineer out of Amazon AWS at a previous company. Whenever one of our cloud services went down, he would go to great lengths to not update our status dashboard. When we finally forced him to update the status page, he would only change it to yellow and write vague updates about how service might be degraded for some customers. He flat out refused to ever admit that the cloud services were down. After some…

At amazon, admitting to a problem will guaranteed lead to having to open a COE, correction of error, which means meetings with executives, inevitable "least effective" rating, development plan, scapegoating, PIP, and firing.

This is complete opposite of my experience at AWS. I'm one of the biggest critics of how we do "software" (just ask any of my managers), but in none of the orgs I've worked at COEs were ever used against you. On the contrary, a good COE is usually applauded.

Re: AWS Cognito is having issues and health dashboards are still green

#209

Earlier quoted context omitted.

By design. If it was a good faith attempt to report status, it would be automatically updated from a flock of canaries instead of through a slow, political process.

Even that would be meaningless at the scale of AWS. "A top of rack switch let out the blue smoke and it'll be ~30 before we can re-rack it" would impact what fraction of a fraction of a percent of canaries? Irrelevant to me, unless of course my VM lives on a box backed by that switch. ;) The status dashboard exists for us to laugh at when things break and to convince C*Os that everything is fine. That's it.

Ehhh... the ratio of "bump in the night" problems that affect just me to genuine outages that cross regions and affect others is about 1:1, and then about 1 in 4 or 5 of the cross-region problems blow up to the scale where they feel forced to update the dashboard. So I disagree, I think a canary flock would be both meaningful and useful.

As you point out, though, the status dashboard isn't truly meant to be either of those things. I don't have any illusions about it ever changing.

Re: AWS Cognito is having issues and health dashboards are still green

#210
Five hours later and nothing has changed. For a company like Amazon this should be unacceptable.

Before someone replies and says use a different AZ, that's not possible for everyone. If you use a 3rd party service that is hosted on us-east-1 you can't do anything about it. For example, many Heroku services are broken because of this.

Post reply on HN