Live data from Hacker News

AWS Cognito is having issues and health dashboards are still green

status.aws.amazon.com

141–150 of 369 posts

Re: AWS Cognito is having issues and health dashboards are still green

#141
post #83

> This is also causing issues with Amplify, API Gateway, AppStream2, AppSync, Athena, Cloudformation, Cloudtrail, Cloudwatch, Cognito, DynamoDB, IoT Services, Lambda, LEX, Managed BlockChain, S3, Sagemaker, and Workspaces. Well, this is a major outgage

We're also seeing issues with FarGate ECS -- the task we had with auto-scaling scaled down to 0. The one we had with a fixed number of workers is fine.

Thanks for mentioning this - that's a nasty failure mode.

Re: AWS Cognito is having issues and health dashboards are still green

#142

Earlier quoted context omitted.

Yes. But I wonder if self-hosting their status page is really the correct decision from a marketing perspective. The people who consumes the status page on say Google Cloud probably know that Google self-hosting it is a bad decision from a technical point of view. So to the only people who care, their choice appear stupid. So I don't really understand what they gain by doing it. I think maybe I am wrong about it bein…

If they host it somewhere else, it signals they lack confidence in their own product. If they self-host it, it signals that they're overconfident in their ability to maintain an accurate status page. Given these two options, which do you think a budget manager will have an easier time signing off on and defending upward?

Yes, that was why I was referring to internal politics and incompetent management.

Re: AWS Cognito is having issues and health dashboards are still green

#144

Many applications – Including Anchor, Adobe Spark, Flickr, SiriusXM and Roku reported disruption caused by this outage. https://news.alphastreet.com/huge-aws-outage-affects-a-wide-...

My iRobot (roomba vacuum robot) app not working for 4 hours...

Re: AWS Cognito is having issues and health dashboards are still green

#145

Earlier quoted context omitted.

Part of the problem is, engineers love shiny things. Status pages are fundamentally boring things. Who wants to work on them? It's always tempting to complicate something simple because in part "ooh shiny", and you can always find reasons to justify why. It takes some strong engineering leadership to effectively argue against complicating things, and not be just a constant pain in the arse to everyone and every thing…

>Part of the problem is, engineers love shiny things. Status pages are fundamentally boring things. Who wants to work on them? I would work on a status page. It's a interesting problem, creating tests that prove services are viable at a place like AWS would be fun. However what I don't want to deal with is some director of so and so I never heard of yelling at me at 3 in the morning because my status page reported th…

Congratulations, you're already complicating the status page.

The status page shouldn't be figuring out what the status of any service is. It's impossible to do without a lot of contextual information about a service and understanding how to evaluate service impact, something that is continually in flux.

It just needs to be a page that is updated manually. AWS has a 24x7 incident management team that could / should do it.

Re: AWS Cognito is having issues and health dashboards are still green

#146

Now is probably a good time to plug some of the open source alternatives to vendor locked in identity solutions: - https://github.com/ory - https://github.com/dexidp/dex - https://github.com/authelia/authelia - https://github.com/keycloak/keycloak - https://www.gluu.org/ - https://github.com/accounts-js/accounts

Fusionauth is pretty cool. I’ve worked with the team a bit on the .net core support.

Re: AWS Cognito is having issues and health dashboards are still green

#147

Earlier quoted context omitted.

That's typical. The AWS status page is a marketing gimmick whose job is to stay green, not a good faith attempt to assess and report status. If there's an outage, seeing it accurately reflected on the status page is the exception, not the rule.

Updating the status dashboard is pretty low priority for operators trying to resolve this issue. It requires escalation up the management chain and careful wording.

By design. If it was a good faith attempt to report status, it would be automatically updated from a flock of canaries instead of through a slow, political process.

Re: AWS Cognito is having issues and health dashboards are still green

#148
post #33

Earlier quoted context omitted.

It is kind of perplexing that AWS dogfoods its own status page. I remember during the massive S3 outage a few years ago that their status page remained green almost the entire time because the red/green/blue icons for the status was stored in... wait for it... S3. You'd think they would have learned from that.

> It is kind of perplexing that AWS dogfoods its own status page. > You'd think they would have learned from that. They did. The page has been updated numerous times since the start of this incident.

it was 1.5 hours before the first service was put on yellow

Re: AWS Cognito is having issues and health dashboards are still green

#149
post #61
post #25

Can anyone explain why status pages are so difficult. Theres even statups like status.io dedicated to this one thing. It really does seem that anytime there is an outage more often than not the status page is showing all green traffic lights. Making it redundant as a tool to corroborate whats happening. How did AWS status page compare with status.io/aws?

> Can anyone explain why status pages are so difficult. What is an outage? When does an outage reach sufficient scale that updating the status page is the right thing to do? I used to work for AWS, and now work for another cloud provider. One thing that's hard to communicate is the sheer scale that these services operate at, what that means architecturally, and how they tend to break. Outages, even just slight degrad…

Posting percentages instead of green/red would fix all of these, no?

Re: AWS Cognito is having issues and health dashboards are still green

#150

Earlier quoted context omitted.

That's typical. The AWS status page is a marketing gimmick whose job is to stay green, not a good faith attempt to assess and report status. If there's an outage, seeing it accurately reflected on the status page is the exception, not the rule.

Updating the status dashboard is pretty low priority for operators trying to resolve this issue. It requires escalation up the management chain and careful wording.

>for operators trying to resolve this issue

It's a shame Amazon doesn't have thousands of employees to divide these tasks between different people, as it is only these busy operators who could update this status page.

If you're right, why have the status page then? It is useless by your definition yes?

Post reply on HN