Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

371–380 of 1001 posts

Re: AWS us-east-1 outage

#371
post #254
post #13

I love that every time this happens, 100% of the services on https://status.aws.amazon.com are green.

It's like trying to get the truth out of a kid that caused some trouble. Mom: Alexa, did you break something? Alexa: No. M: Really? What's this? 500 Internal server error A: ok maybe management console is down M: Anything else? A: ... A: ... ok maybe cloudwatch logs M: Ah hah. What else? A: That's it, I swear! M: 503 ClientError A: ...well okay secretsmanager might be busted too...

There was a great response in r/relationship advice the other day where someone said that OP's partner forced a fight because they're planning to cheat on them, reconcile, and then will 'trickle out the truth' over the next 6 months. I'm stealing that phrase.

Re: AWS us-east-1 outage

#372
post #245

Looks like they've acknowledged it on the status page now. https://status.aws.amazon.com/ > 8:22 AM PST We are investigating increased error rates for the AWS Management Console. > 8:26 AM PST We are experiencing API and console issues in the US-EAST-1 Region. We have identified root cause and we are actively working towards recovery. This issue is affecting the global console landing page, which is also hosted in US…

Yeah, but I still have a different understanding what "Increased Error Rates" means.

IMHO it should mean that the rate of errors is increased but the service is still able to serve a substantial amount of traffic. If the rate of errors is bigger than, let's say, 90% that's not an increased error rate, that's an outage.

Re: AWS us-east-1 outage

#373

Earlier quoted context omitted.

One time gcp argued that since they did return 404s on gcs for a few hours that wasn’t an uptime/latency sla violation so we were not entitled to refund (tho they refunded us anyway)

Man, between costs and shenanigans like this, why don't more companies self-host?

1. Leadership prefers to blame cloud when things break rather than take responsibility.

2. Cost is not an issue (until it is but you’re already locked in so oh well)

3. Faang has drained the talent pool of people who know how

Re: AWS us-east-1 outage

#374
post #254

Earlier quoted context omitted.

It's like trying to get the truth out of a kid that caused some trouble. Mom: Alexa, did you break something? Alexa: No. M: Really? What's this? 500 Internal server error A: ok maybe management console is down M: Anything else? A: ... A: ... ok maybe cloudwatch logs M: Ah hah. What else? A: That's it, I swear! M: 503 ClientError A: ...well okay secretsmanager might be busted too...

Funny I literally just asked my Alexa. Me: Alexa, is AWS down right now? Alexa: I'd rather not answer that

Wise robot.

That's a bit like involving your kid in an argument between parents.

Re: AWS us-east-1 outage

#375
post #62

Earlier quoted context omitted.

They can do this without an alliance. They very intentionally choose not to do it. Every major company has moved away from having accurate status pages.

It's because none of these companies are held responsible for missing their actual SLAs, as opposed to their self-reported SLA compliance. So unless regulation gets implemented that says otherwise, there's zero incentive for any company to maintain an accurate status page.

How did you find a way to bring regulations into this? There are monitoring services you can pay for to keep an eye on your SLAs and your vendors'.

If not happy with the results switch.

Re: AWS us-east-1 outage

#376
post #66

Are the actual services down, or is it just the console and/or login page? For example, the sign-up page appears to be working: https://portal.aws.amazon.com/billing/signup#/start Are websites that run on AWS us-east up? Are the AWS CLIs working?

One of my sites went offline an hour ago because the web server stopped responding. I can't SSH into it or get any type of response. The database server in the same region and zone is continuing to run fine though.

Interesting, is the site on a particular type of EC2 instance, e.g. bare metal? I see c4.xlarge is doing fine in us-east-1.

Re: AWS us-east-1 outage

#377

Earlier quoted context omitted.

I don't think engineers can believe in no-blame analysis if they know it'll harm career growth. I can't unilaterally promote John Doe, I have to convince other leaders that John would do well the next level up. And in those discussions, they could bring up "but John has caused 3 incidents this year", and honestly, maybe they'd be right.

Would they? Having 3 outages in a year sounds like an organization problem. Not enough safeguards to prevent very routine human errors. But instead of worrying about that we just assign a guy to take the fall

Well if John caused 3 outages and and his peers Sally and Mike each caused 0, it's worth taking a deeper look. There's a real possibility he's getting screwed by a messed up org, also he could be doing slapdash work or he seriously might not undertsand the seriousness of an outage.

Re: AWS us-east-1 outage

#378

Earlier quoted context omitted.

I don't think engineers can believe in no-blame analysis if they know it'll harm career growth. I can't unilaterally promote John Doe, I have to convince other leaders that John would do well the next level up. And in those discussions, they could bring up "but John has caused 3 incidents this year", and honestly, maybe they'd be right.

Would they? Having 3 outages in a year sounds like an organization problem. Not enough safeguards to prevent very routine human errors. But instead of worrying about that we just assign a guy to take the fall

If you work in a technical role and you _don't_ have the ability to break something, you're unlikely to be contributing in a significant way. Likely that would make you a junior developer whose every line of code is heavily scrutinized.

Engineers should be experts and you should be able to trust them to make reasonable choices about the management of their projects.

That doesn't mean there can't be some checks in place, and it doesn't mean that all engineers should be perfect.

But you also have to acknowledge that adding all of those safeties has a cost. You can be a competent person who requires fewer safeties or less competent with more safeties.

Which one provides more value to an organization?

Re: AWS us-east-1 outage

#379
post #113

Friends tell friends to pick us-east-2. Virginia is for lovers, Ohio is for availability.

If you're not multi-cloud in 2021 and are expecting 5-9's, I feel bad for you.

I imagine there are very few businesses where the extra cost of going multi-cloud is smaller than the cost of being down during AWS outages.

Re: AWS us-east-1 outage

#380

Earlier quoted context omitted.

EC2 or S3 showing red in any region literally requires personal approval of the CEO of AWS.

Is this true or a joke? This sort of policy is how you destroy trust.

If you trust them at this point, you have not being paying attention, and will probably continue to trust after this.
Post reply on HN