Live data from Hacker News

AWS Cognito is having issues and health dashboards are still green

status.aws.amazon.com

101–110 of 369 posts

Re: AWS Cognito is having issues and health dashboards are still green

#101

Earlier quoted context omitted.

When your company gets sufficiently large, outages become political. Failure happens at the speed of computing but agreeing that something is failing in a way that customers need to be told about is a slower process. Even when status pages are fully automatic (rather than manually updated), there will tend to be gaming of the metrics that constitute that. Ideally you would just be monitoring your SLOs and publishing…

And not just outages, but security incidents. I’ve worked at/with/for many companies as both an employee and a consultant where the top priority wasn’t to have fewer security incidents, but to have fewer security incidents that would require disclosure . Publicly disclosing an incident to a customer is embarrassing and potentially damaging but almost equally as damaging is telling other teams you had an incident. Now…

All of those are reasons that the determination of status should be totally independent of the company technically and legally.

Re: AWS Cognito is having issues and health dashboards are still green

#102
post #67
post #53

Earlier quoted context omitted.

K so they avoided that problem, but something similar has obviously gone wrong again, considering that Kinesis had been partially or fully down for almost an hour before the status page got their first update. And the fact remains that currently an outage of AWS's own infrastructure is impacting AWS's ability to status updates on its own status dashboard. It's just seems so... amateurish.

That's incredibly annoying, given the mandate the replacement service had. I'd be curious to be a fly on the wall during the next Ops meeting when it comes up that yet again the status dashboard got made in a way that makes it hard to update during an outage.

Maybe they should ask a question about resilient status page architecture among the ridiculous coding riddles they give candidates...lol!

Re: AWS Cognito is having issues and health dashboards are still green

#105

> This is also causing issues with Amplify, API Gateway, AppStream2, AppSync, Athena, Cloudformation, Cloudtrail, Cloudwatch, Cognito, DynamoDB, IoT Services, Lambda, LEX, Managed BlockChain, S3, Sagemaker, and Workspaces. Well, this is a major outgage

yep, iot in us-east-1 not working for me

Re: AWS Cognito is having issues and health dashboards are still green

#106

> This is also causing issues with Amplify, API Gateway, AppStream2, AppSync, Athena, Cloudformation, Cloudtrail, Cloudwatch, Cognito, DynamoDB, IoT Services, Lambda, LEX, Managed BlockChain, S3, Sagemaker, and Workspaces. Well, this is a major outgage

Indeed, we had the first AWS Kinesis issues already at 13:50 (UTC). Now it's still ongoing after two hours. The status page didn't even update in the first 45 min or so...

That's typical. The AWS status page is a marketing gimmick whose job is to stay green, not a good faith attempt to assess and report status. If there's an outage, seeing it accurately reflected on the status page is the exception, not the rule.

Re: AWS Cognito is having issues and health dashboards are still green

#107
post #94

Earlier quoted context omitted.

It's common practice for small players but Amazon, Microsoft Azure and Google Cloud host their status pages on their own servers because they value the marketing aspect higher than a functioning status page for their customers.

I find it surprising how many people forget how much underlying business motives drive pretty much every action they make and how this is quickly forgotten by many. No matter how much you value science and engineering, it ultimately doesn't matter to the business unless that aligns directly with their revenue stream. Sometimes it does, sometimes it doesn't.

Yes. But I wonder if self-hosting their status page is really the correct decision from a marketing perspective. The people who consumes the status page on say Google Cloud probably know that Google self-hosting it is a bad decision from a technical point of view. So to the only people who care, their choice appear stupid.

So I don't really understand what they gain by doing it. I think maybe I am wrong about it being a marketing concern and that the choice is more related to internal politics and incompetent management.

Re: AWS Cognito is having issues and health dashboards are still green

#108

Cognito is one of the most frustrating AWS services I have to work with, it is almost, but not quite, entirely unlike an SP. We're using it to federate customer IDPs through user pools, but this ends up with customer configs being region specific. Has anyone figured out how to set up Cognito in multiple regions without the hijinx of having the customer setup trusts for each region? Not to mention, while multiple trus…

Eh? Brokering amongst multiple trusts (and managing protocol transition) is almost the raison d'etre for lifting token issuance out of your app and into ADFS, Okta, Auth0, etc.

Of course you'll have to deal with home realm discovery--really need to go in with open eyes on that one.

Re: AWS Cognito is having issues and health dashboards are still green

#109
post #81

"I want to have an AWS region where everything breaks with high frequency..."[0] discussed here [1] [0] https://twitter.com/apgwoz/status/1292519906433306625?s=20 [1] https://news.ycombinator.com/item?id=24103746

Isn't that just called us-east-1?

I've read this multiple times that AWS us-east-1 region is the one that has the highest number of outages. I am eager to hear others' experiences here.
Post reply on HN