Live data from Hacker News

AWS Cognito is having issues and health dashboards are still green

status.aws.amazon.com

91–100 of 369 posts

Re: AWS Cognito is having issues and health dashboards are still green

#91
post #74

Earlier quoted context omitted.

Yes but only if you initiate a claim and follow their steps. Check out these onerous terms: Credit Request and Payment Procedures To receive a Service Credit, you must submit a claim by opening a case in the AWS Support Center. To be eligible, the credit request must be received by us by the end of the second billing cycle after which the incident occurred and must include: 1. the words “SLA Credit Request” in the su…

> Yes but only if you initiate a claim and follow their steps. Check out these onerous terms: There's most likely a reason for this. Like, maybe in the past AWS customers have tried claiming for SLA credits for incidents that didn't impact them, in order to reduce their bill.

This is backwards thinking. Why require customers to file a claim for what are obvious outages? Instead, AWS should automatically apply credits to those accounts that have paid for guaranteed uptime without requiring this whole silly claims process.

The mechanism can be really simple. If AWS themselves posts an outage to their status page and/or some third-party service posts an outage then credits are immediately applied to the services where there are outages for those that paid for high level uptime guarantees without requiring any claims process. It can easily be done if they want to do it that way.

Of course from a business perspective I understand why they're doing it the way that they are. If they can make customers jump through hoops, then only those who really care will follow through. Meanwhile the uptime guarantee can continue as an empty promise.

Re: AWS Cognito is having issues and health dashboards are still green

#92

"I want to have an AWS region where everything breaks with high frequency..."[0] discussed here [1] [0] https://twitter.com/apgwoz/status/1292519906433306625?s=20 [1] https://news.ycombinator.com/item?id=24103746

`us-wtf-1`

Re: AWS Cognito is having issues and health dashboards are still green

#93
post #60

Any tips on how to collect on SLA credits from this?

The procedure is outlined in each service's SLA -- though I think they're all pretty much the same.

Annoyingly, they expect you to do the leg work to show when the outage happened and supply logs demonstrating that you were impacted.

Might want to do some napkin math first to see if the amount credit is worth your time. The couple times my org considered pursuing it, it just wasn't worth the effort. (Though, personally, I think that speaks to a larger problem with the SLA.)

Credit Request Procedure in Kinesis SLA: https://aws.amazon.com/kinesis/sla/#Credit_Request_and_Payme...

Re: AWS Cognito is having issues and health dashboards are still green

#94

Isn't it common practice to host your status board on someone else's infrastructure? In 2017 there was an S3 issue that supposedly affected their ability to post. I believe they said that they were updating how they posted to the status board so that there would no longer be a dependency on S3. Well, I guess whatever they're dependent on now broke.

It's common practice for small players but Amazon, Microsoft Azure and Google Cloud host their status pages on their own servers because they value the marketing aspect higher than a functioning status page for their customers.

I find it surprising how many people forget how much underlying business motives drive pretty much every action they make and how this is quickly forgotten by many.

No matter how much you value science and engineering, it ultimately doesn't matter to the business unless that aligns directly with their revenue stream. Sometimes it does, sometimes it doesn't.

Re: AWS Cognito is having issues and health dashboards are still green

#95

"I want to have an AWS region where everything breaks with high frequency..."[0] discussed here [1] [0] https://twitter.com/apgwoz/status/1292519906433306625?s=20 [1] https://news.ycombinator.com/item?id=24103746

`us-wtf-1`

ive... never loved a region before

Re: AWS Cognito is having issues and health dashboards are still green

#97
post #88
post #31

Earlier quoted context omitted.

You've missed the point entirely, bravo

Ughh, he claims to have been scammed by AWS because their services are having an outage, and I'm missing the point? This stupid hyperboles need to be shot down. I'm sick and tired of the victim mentality and hyperboles. Every time something inconvenient happens, people scream and shout at the top of their lungs like the world has wronged them. NO, YOU DID NOT GET SCAMMED BY AMAZON BECAUSE THEY HAVE A SERVICE OUTAGE.…

> NO, YOU DID NOT GET SCAMMED BY AMAZON BECAUSE THEY HAVE A SERVICE OUTAGE.

Many people have pushed for cloud services because they are supposed to be more reliable than setting up a system in a rack in a datacenter. AWS will constantly point to their uptime guarantee, except it isn't a guarantee. It's just a sales tactic that misrepresents the historical uptime of AWS.

The larger point is that if 99.5% is the real expected uptime, it's vastly cheaper to have a solution that is not AWS, even before you factor in the cost savings of learning their security model and completely opaque billing system.

Advertising a product with features it does not have is the classic definition of a scam.

Re: AWS Cognito is having issues and health dashboards are still green

#98

Earlier quoted context omitted.

When your company gets sufficiently large, outages become political. Failure happens at the speed of computing but agreeing that something is failing in a way that customers need to be told about is a slower process. Even when status pages are fully automatic (rather than manually updated), there will tend to be gaming of the metrics that constitute that. Ideally you would just be monitoring your SLOs and publishing…

And not just outages, but security incidents. I’ve worked at/with/for many companies as both an employee and a consultant where the top priority wasn’t to have fewer security incidents, but to have fewer security incidents that would require disclosure . Publicly disclosing an incident to a customer is embarrassing and potentially damaging but almost equally as damaging is telling other teams you had an incident. Now…

Additionally, you're penalized for doing it "right", because you're often competing against companies which rarely say that anything's wrong (ahem, Mailchimp). You look worse, because you're being transparent about service status, which creates the perception that you're generally less stable.

Re: AWS Cognito is having issues and health dashboards are still green

#99
post #57
post #22

There's a lot more going on over there... - 7 cloudfront distributions created today are still in "InProgress", a few already for more than one hour - The support case I created about it doesn't show up in my support portal. Direct link to it does work though

I think the issue is that Kinesis is a single point of failure for a ton of systems. When it goes down, loads of other system's workflows can't operate. AWS is famous for eating their own dog food and someone just poisoned it.

Maybe they bought the dogfood from the Amazon Marketplace, but it was counterfeit.

Re: AWS Cognito is having issues and health dashboards are still green

#100
post #73

Earlier quoted context omitted.

That day was a nightmare for a lot of people - it wasn't just S3 that went down, it was like all of US-EAST. Luckily my company decided against multi-az for the cost savings so I spent all day firefighting.

So what’s the cost breakdown? Did they make the right decision?

For one day of his time and probably a small part of a day of diminished service, most likely.
Post reply on HN