Live data from Hacker News

AWS Cognito is having issues and health dashboards are still green

status.aws.amazon.com

251–260 of 369 posts

Re: AWS Cognito is having issues and health dashboards are still green

#251

We hired an engineer out of Amazon AWS at a previous company. Whenever one of our cloud services went down, he would go to great lengths to not update our status dashboard. When we finally forced him to update the status page, he would only change it to yellow and write vague updates about how service might be degraded for some customers. He flat out refused to ever admit that the cloud services were down. After some…

That's opposite of my experience at AWS. It's likely that the culture at AWS has changed over the past few years, it's also likely that there's a difference in culture between teams.

I have no doubt it varies from team to team. Like I said, my other friends at Amazon had more positive experiences.

I assume there's some selection bias going on whenever we're able to hire people out of FAANG companies. We compensated similarly, but in theory had a lower promotion ceiling simply because we weren't FAANG. I assume he wanted out of Amazon because he wasn't on a great team there.

Re: AWS Cognito is having issues and health dashboards are still green

#252

We hired an engineer out of Amazon AWS at a previous company. Whenever one of our cloud services went down, he would go to great lengths to not update our status dashboard. When we finally forced him to update the status page, he would only change it to yellow and write vague updates about how service might be degraded for some customers. He flat out refused to ever admit that the cloud services were down. After some…

At amazon, admitting to a problem will guaranteed lead to having to open a COE, correction of error, which means meetings with executives, inevitable "least effective" rating, development plan, scapegoating, PIP, and firing.

What the hell is a COE? I hate that nobody seems to bother defining their acronyms anymore.

Re: AWS Cognito is having issues and health dashboards are still green

#253
post #74

Earlier quoted context omitted.

Yes but only if you initiate a claim and follow their steps. Check out these onerous terms: Credit Request and Payment Procedures To receive a Service Credit, you must submit a claim by opening a case in the AWS Support Center. To be eligible, the credit request must be received by us by the end of the second billing cycle after which the incident occurred and must include: 1. the words “SLA Credit Request” in the su…

> Yes but only if you initiate a claim and follow their steps. Check out these onerous terms: There's most likely a reason for this. Like, maybe in the past AWS customers have tried claiming for SLA credits for incidents that didn't impact them, in order to reduce their bill.

When my ISP was unable to provide connectivity for an extended period, it automatically compensated me. I didn't have to do anything. The relevant system was being monitored, the ISP knew exactly when it was out of service, and I was credited accordingly with an apology and a note on my next bill showing the reduction. It doesn't seem unreasonable to expect the biggest name in the cloud to do something similar to support its customers when it screws up.

Re: AWS Cognito is having issues and health dashboards are still green

#254
post #237

Earlier quoted context omitted.

Literally 99.9% of the employees have no more knowledge than you about the inner workings of a given AWS service. This isn't to forgive their lack of updating the status page, but large engineering orgs are never the knowledge monoliths you might imagine they are.

This isn't a question of knowing a service is down. We're assuming the team that is fixing the service, knows it is down. It was a question of not having the resources to direct literally any other person in the org to log into an admin panel and flip a toggle from green to red.

I was merely addressing the "thousands of employees" non sequitur. Org structure means that the raw number of employees is a meaningless metric. The only people who are going to potentially flip that switch are going to have some sort of direct responsibility for the product. That number is going to be very similar whether it's a large company like Amazon or a smaller one like, say, Heroku or Dreamhost.

Re: AWS Cognito is having issues and health dashboards are still green

#255
post #28
post #11

Earlier quoted context omitted.

This has happened a few times before, actually. Dogfooding is good, but not for status pages! Cloudflare does it right for their status page ( https://www.cloudflarestatus.com ). They don't use Cloudflare itself for it (you can tell because /cdn-cgi/trace returns nothing), the actual backend is Atlassian Statuspage, their TLS certificate is issued by Let's Encrypt instead of Cloudflare itself, and it's on a completel…

Have you checked https://www.githubstatus.com/ ? ;-)

GitHub doesn't own their own datacenters.

Re: AWS Cognito is having issues and health dashboards are still green

#256
post #221

Earlier quoted context omitted.

Yeah, had the same experience at a previous company. It's very frustrating that your transparency gets used against you by unscrupulous competitors.

How is it unscrupulous? This sort of shit happens all the time at all levels. Companies use each other’s public specs in their competition all the time. Or capitalizing on features like headphone jacks etc. in their ads before proceeding to remove them from their own products anyway (Samsung and Google) and so on.

You're missing the point. The point is that it isn't apple to apples. If you are honest with a dashboard, and the competitor isn't (or doesn't have one), it's not fair to compare.

Re: AWS Cognito is having issues and health dashboards are still green

#257

We hired an engineer out of Amazon AWS at a previous company. Whenever one of our cloud services went down, he would go to great lengths to not update our status dashboard. When we finally forced him to update the status page, he would only change it to yellow and write vague updates about how service might be degraded for some customers. He flat out refused to ever admit that the cloud services were down. After some…

No idea what happens on AWS as I don't work there, but I have another perspective on this. There are perverse incentives to NOT update your status dashboard. Once I was asked by management to _take our status dashboard down_ . That sounded backwards, so I dug a bit more. Turns out our competitor was using our status dashboard as ammo against us in their sales pitch. Their claim was that we had too many issues and wer…

I'd monitor the competition and use it to your advantage.

Re: AWS Cognito is having issues and health dashboards are still green

#258
> 2:43 PM PST Between 5:15 AM and 2:28 PM PST customers experienced increased API failure rates for Cognito User Pools and Identity Pools in the US-EAST-1 Region. This was due to an issue with Kinesis Data Streams. We have implemented a mitigation to this issue. Cognito is now operating normally.

Seems like they fixed Cognito while Kinesis and many other services are still broken - presumably somehow removing the dependency on Kinesis? It’ll be really interesting if their post mortem explains this mitigation.

Re: AWS Cognito is having issues and health dashboards are still green

#259

We hired an engineer out of Amazon AWS at a previous company. Whenever one of our cloud services went down, he would go to great lengths to not update our status dashboard. When we finally forced him to update the status page, he would only change it to yellow and write vague updates about how service might be degraded for some customers. He flat out refused to ever admit that the cloud services were down. After some…

> FWIW, I have other friends who work on different teams at Amazon who have not had such bad experiences.

Right, "it depends on the team!!" as we always hear whenever this stuff comes up. But there should be zero teams doing to people what your engineer's AWS team did to him.

Re: AWS Cognito is having issues and health dashboards are still green

#260
post #195

Earlier quoted context omitted.

Have worked at AWS before, and I can attest to this. Whenever we had an outage, our director and senior manager would take a call on whether to update the dashboard or not. Having 'red' dashboard catches lot of eyes, so people responsible for making this decision always look at it from political point of view. As a dev oncall, we used to get 20 sev2s per day (an oncall ticket which needs to be handled within 15 mins)…

Wow. If I were in charge, the team running a service should not be the same team who decides whether a given service is healthy. This is pretty damaging info about the unprofessional way AWS actually appears to be run.

I think this is pretty typical, as often outsiders don't have the visibility into the issue to determine whether there's an issue.
Post reply on HN