Live data from Hacker News

AWS Cognito is having issues and health dashboards are still green

status.aws.amazon.com

281–290 of 369 posts

Re: AWS Cognito is having issues and health dashboards are still green

#281

Earlier quoted context omitted.

Congratulations, you're already complicating the status page. The status page shouldn't be figuring out what the status of any service is. It's impossible to do without a lot of contextual information about a service and understanding how to evaluate service impact, something that is continually in flux. It just needs to be a page that is updated manually. AWS has a 24x7 incident management team that could / should d…

Updated manually by whom? I'm afraid you're shifting the complexity to a manual process. I agree that it doesn't have to, and perhaps should not, be fully automated. But automating some parts will help not waste time on last minute arguments.

> Updated manually by whom? > I'm afraid you're shifting the complexity to a manual process.

You're right, that's 100% what I'm doing. Why? Because it shouldn't be that complicated to update an overall health status page during an outage event, and it shouldn't take other tools and services within AWS to do it.

A common pattern in cloud providers (including AWS) is that services have some kind of tiering, whereby you can't pick up a dependency on any service on a lower tier than yourselves. Tier 2 services can't rely on Tier 3 services, etc. Services like, say, IAM, would be right at the very top. It can't rely on EBS, ELB etc. Everything has to be created in-service, because everything ultimately has to rely on authentication working.

If they're going to keep an overall status page going, it needs to be seen as a top tier service, just like identity is. That's where they were headed towards when I left AWS about 5 1/2 years ago. It had been spurred by a previous major incident couldn't be reflected in the status dashboard because of a failure in a dependency.

> I agree that it doesn't have to, and perhaps should not, be fully automated. But automating some parts will help not waste time on last minute arguments.

I go in to a bit more detail in another comment within this discussion, but a status page does not even close to accurately capture the ways that cloud environments fail, which are very, very rarely affecting more than a small percentage of customers, and even then often in some very specific way under specific circumstances. That's why AWS built the personalised status page service. They want to ensure that customers have an accurate way of telling what is going on with services they're consuming, rather than the confusing situation of checking an overall status site that doesn't really reflect their experience and never could.

Situations like today's where it at least (from the outside) seemed like Kinesis was completely down, would be a good example of something that should be reflect in the main overall status page.

The status page should be manual, and should be something the incident management team can do (and have political ability to force it to happen, rather than being subject to the whims of service directors)

Re: AWS Cognito is having issues and health dashboards are still green

#282

Earlier quoted context omitted.

Amazon was at least aware enough to recognized that AWS circular dependencies were a bad thing. From what I heard they had to make changes. A big problem is the largest services like S3. If part of S3 were to use DynamoDB and DynamoDB used S3, then if one goes down, they might never restart either service. There is strong manager incentive at Amazon to build on other services as a way to ingratiate with other manager…

Fascinating, I hadn't even considered how the org design and incentives in place internally at AWS affects the way some of the outward facing services are designed. Is an example say an up and coming director wanted to build a new service that depends on an existing service to curry favor? Do you have more examples or anecdotes to share?

Isn’t this just an example of Conway’s law? https://en.m.wikipedia.org/wiki/Conway%27s_law

Re: AWS Cognito is having issues and health dashboards are still green

#283

Earlier quoted context omitted.

Reminds me of a recent outage from IBM Cloud, where the VPN was hosted on IBM Cloud so employees couldn’t log in to fix it, and the email is hosted on IBM Cloud so support teams couldn’t email customers to let them know and even access to their Twitter was behind the non-functional VPN so they couldn’t tweet during the outage either.

This sounds like a true nightmare. Do you have a link for more details?

I actually heard the details on a cloud-related podcast, but I found a transcript of the episode here: https://www.lastweekinaws.com/podcast/aws-morning-brief/whit...

The relevant bit:

>[customers] were texting with their account managers, because the account managers had no access to any internal systems. Reportedly, the corporate VPN was not working. My thesis is... everything was single-tracking through a corporate VPN that itself was subject to this disruption... their traditional tweets have been done through an enterprise social media client called Sprinklr

Re: AWS Cognito is having issues and health dashboards are still green

#284
post #195

We hired an engineer out of Amazon AWS at a previous company. Whenever one of our cloud services went down, he would go to great lengths to not update our status dashboard. When we finally forced him to update the status page, he would only change it to yellow and write vague updates about how service might be degraded for some customers. He flat out refused to ever admit that the cloud services were down. After some…

Have worked at AWS before, and I can attest to this. Whenever we had an outage, our director and senior manager would take a call on whether to update the dashboard or not. Having 'red' dashboard catches lot of eyes, so people responsible for making this decision always look at it from political point of view. As a dev oncall, we used to get 20 sev2s per day (an oncall ticket which needs to be handled within 15 mins)…

If you don't get the non-existant relation between Sev2s and the public dashboard, then that says a lot about your qualification to post about AWS on HN.

Re: AWS Cognito is having issues and health dashboards are still green

#285

Earlier quoted context omitted.

PIP : Performance Improvement Plan I assume this to mean that it is an element of an individual's PIP, a formal process to set guides for getting someone to commit to a higher level of achievement.

While on its face PIP is a guide to getting someone to commit to a higher level of improvement, for many companies its a formal warning that you need to shape up or you're going to be let go.

> for many companies its a formal warning that you need to shape up or you're going to be let go.

Patently incorrect. A PIP is management telling you that you need to seek alternative employment, now.

Joking/sarcasm aside: I’ve never seen or heard someone who is placed on a PIP successfully “exit” the PIP. They exit the company or they’re exited from the company. PIPs seem to mark the start of the “we are building formal documentation to fire you” phase of losing a job.

Re: AWS Cognito is having issues and health dashboards are still green

#286
post #221

Earlier quoted context omitted.

Yeah, had the same experience at a previous company. It's very frustrating that your transparency gets used against you by unscrupulous competitors.

How is it unscrupulous? This sort of shit happens all the time at all levels. Companies use each other’s public specs in their competition all the time. Or capitalizing on features like headphone jacks etc. in their ads before proceeding to remove them from their own products anyway (Samsung and Google) and so on.

OK, so what's your point? The outcome of this is still a worse situation for everyone involved in the end.

Re: AWS Cognito is having issues and health dashboards are still green

#287

Earlier quoted context omitted.

Wow. If I were in charge, the team running a service should not be the same team who decides whether a given service is healthy. This is pretty damaging info about the unprofessional way AWS actually appears to be run.

It's funny that you point to that as the problem. The problem is more AWS' toxic engineering culture that has engineers fearing for their jobs in a way that guides their decision making. It's bad company culture, end of story.

AWS is big. Amazon is even bigger. Disgruntled people are the ones who often cry the loudest. Just because there may be teams who act like this, doesn't mean that is the case in general.

You don't hear a lot of people praising AWS, the same way you don't hear a lot of people saying how great it is to have an iPhone. If I am happy, I have little incentive to post about it, since that should be the default state.

But the matter of fact is simple. If you end up in a team like this, switch and raise complaints afterwards. Nothing stops you from it. There is no "toxic engineering culture" at AWS. The problem is that AWS makes you into an owner and that includes owning your career. That means if you feel something is wrong, YOU are expected to act. No one will do it for you. And there are plenty of mechanism for you to act.

This is the greatest benefit of working at Amazon but its also the downfall of people who are not able to own things.

Re: AWS Cognito is having issues and health dashboards are still green

#288

Earlier quoted context omitted.

While on its face PIP is a guide to getting someone to commit to a higher level of improvement, for many companies its a formal warning that you need to shape up or you're going to be let go.

> for many companies its a formal warning that you need to shape up or you're going to be let go. Patently incorrect. A PIP is management telling you that you need to seek alternative employment, now . Joking/sarcasm aside: I’ve never seen or heard someone who is placed on a PIP successfully “exit” the PIP. They exit the company or they’re exited from the company. PIPs seem to mark the start of the “we are building f…

ya i agree.

Re: AWS Cognito is having issues and health dashboards are still green

#289

Earlier quoted context omitted.

It's funny that you point to that as the problem. The problem is more AWS' toxic engineering culture that has engineers fearing for their jobs in a way that guides their decision making. It's bad company culture, end of story.

AWS is big. Amazon is even bigger. Disgruntled people are the ones who often cry the loudest. Just because there may be teams who act like this, doesn't mean that is the case in general. You don't hear a lot of people praising AWS, the same way you don't hear a lot of people saying how great it is to have an iPhone. If I am happy, I have little incentive to post about it, since that should be the default state. But t…

> The problem is that AWS makes you into an owner and that includes owning your career.

Firing me for correctly telling customers that their services are down is not my idea of making me an owner.

Re: AWS Cognito is having issues and health dashboards are still green

#290

Earlier quoted context omitted.

What is PIP?

Performance Improvement Plan, they are not unique to Amazon, most places have them though the process may differ. Not to be too cynical but ultimately they’re a way to document that you’re not meeting expectations - before being fired. Should there be any sort of employment claim later its a mechanism by which an employer can show documentation that any issues related to your being let go were performance related and…

I think the reason they don’t work is because someone doesn’t just magically become a better employee over two months.
Post reply on HN