Earlier quoted context omitted.
When your company gets sufficiently large, outages become political. Failure happens at the speed of computing but agreeing that something is failing in a way that customers need to be told about is a slower process. Even when status pages are fully automatic (rather than manually updated), there will tend to be gaming of the metrics that constitute that. Ideally you would just be monitoring your SLOs and publishing…
And not just outages, but security incidents. I’ve worked at/with/for many companies as both an employee and a consultant where the top priority wasn’t to have fewer security incidents, but to have fewer security incidents that would require disclosure . Publicly disclosing an incident to a customer is embarrassing and potentially damaging but almost equally as damaging is telling other teams you had an incident. Now…
AWS Cognito is having issues and health dashboards are still green
101–110 of 369 posts
Re: AWS Cognito is having issues and health dashboards are still green
#102Earlier quoted context omitted.
K so they avoided that problem, but something similar has obviously gone wrong again, considering that Kinesis had been partially or fully down for almost an hour before the status page got their first update. And the fact remains that currently an outage of AWS's own infrastructure is impacting AWS's ability to status updates on its own status dashboard. It's just seems so... amateurish.
That's incredibly annoying, given the mandate the replacement service had. I'd be curious to be a fly on the wall during the next Ops meeting when it comes up that yet again the status dashboard got made in a way that makes it hard to update during an outage.
Re: AWS Cognito is having issues and health dashboards are still green
#103Re: AWS Cognito is having issues and health dashboards are still green
#104Re: AWS Cognito is having issues and health dashboards are still green
#105> This is also causing issues with Amplify, API Gateway, AppStream2, AppSync, Athena, Cloudformation, Cloudtrail, Cloudwatch, Cognito, DynamoDB, IoT Services, Lambda, LEX, Managed BlockChain, S3, Sagemaker, and Workspaces. Well, this is a major outgage
Re: AWS Cognito is having issues and health dashboards are still green
#106> This is also causing issues with Amplify, API Gateway, AppStream2, AppSync, Athena, Cloudformation, Cloudtrail, Cloudwatch, Cognito, DynamoDB, IoT Services, Lambda, LEX, Managed BlockChain, S3, Sagemaker, and Workspaces. Well, this is a major outgage
Indeed, we had the first AWS Kinesis issues already at 13:50 (UTC). Now it's still ongoing after two hours. The status page didn't even update in the first 45 min or so...
Re: AWS Cognito is having issues and health dashboards are still green
#107Earlier quoted context omitted.
It's common practice for small players but Amazon, Microsoft Azure and Google Cloud host their status pages on their own servers because they value the marketing aspect higher than a functioning status page for their customers.
I find it surprising how many people forget how much underlying business motives drive pretty much every action they make and how this is quickly forgotten by many. No matter how much you value science and engineering, it ultimately doesn't matter to the business unless that aligns directly with their revenue stream. Sometimes it does, sometimes it doesn't.
So I don't really understand what they gain by doing it. I think maybe I am wrong about it being a marketing concern and that the choice is more related to internal politics and incompetent management.
Re: AWS Cognito is having issues and health dashboards are still green
#108Cognito is one of the most frustrating AWS services I have to work with, it is almost, but not quite, entirely unlike an SP. We're using it to federate customer IDPs through user pools, but this ends up with customer configs being region specific. Has anyone figured out how to set up Cognito in multiple regions without the hijinx of having the customer setup trusts for each region? Not to mention, while multiple trus…
Of course you'll have to deal with home realm discovery--really need to go in with open eyes on that one.
Re: AWS Cognito is having issues and health dashboards are still green
#109"I want to have an AWS region where everything breaks with high frequency..."[0] discussed here [1] [0] https://twitter.com/apgwoz/status/1292519906433306625?s=20 [1] https://news.ycombinator.com/item?id=24103746
Isn't that just called us-east-1?