I love that every time this happens, 100% of the services on https://status.aws.amazon.com are green.
It's like trying to get the truth out of a kid that caused some trouble. Mom: Alexa, did you break something? Alexa: No. M: Really? What's this? 500 Internal server error A: ok maybe management console is down M: Anything else? A: ... A: ... ok maybe cloudwatch logs M: Ah hah. What else? A: That's it, I swear! M: 503 ClientError A: ...well okay secretsmanager might be busted too...
AWS us-east-1 outage
371–380 of 1001 posts
Re: AWS us-east-1 outage
#372Looks like they've acknowledged it on the status page now. https://status.aws.amazon.com/ > 8:22 AM PST We are investigating increased error rates for the AWS Management Console. > 8:26 AM PST We are experiencing API and console issues in the US-EAST-1 Region. We have identified root cause and we are actively working towards recovery. This issue is affecting the global console landing page, which is also hosted in US…
IMHO it should mean that the rate of errors is increased but the service is still able to serve a substantial amount of traffic. If the rate of errors is bigger than, let's say, 90% that's not an increased error rate, that's an outage.
Re: AWS us-east-1 outage
#373Earlier quoted context omitted.
One time gcp argued that since they did return 404s on gcs for a few hours that wasn’t an uptime/latency sla violation so we were not entitled to refund (tho they refunded us anyway)
Man, between costs and shenanigans like this, why don't more companies self-host?
2. Cost is not an issue (until it is but you’re already locked in so oh well)
3. Faang has drained the talent pool of people who know how
Re: AWS us-east-1 outage
#374Earlier quoted context omitted.
It's like trying to get the truth out of a kid that caused some trouble. Mom: Alexa, did you break something? Alexa: No. M: Really? What's this? 500 Internal server error A: ok maybe management console is down M: Anything else? A: ... A: ... ok maybe cloudwatch logs M: Ah hah. What else? A: That's it, I swear! M: 503 ClientError A: ...well okay secretsmanager might be busted too...
Funny I literally just asked my Alexa. Me: Alexa, is AWS down right now? Alexa: I'd rather not answer that
That's a bit like involving your kid in an argument between parents.
Re: AWS us-east-1 outage
#375Earlier quoted context omitted.
They can do this without an alliance. They very intentionally choose not to do it. Every major company has moved away from having accurate status pages.
It's because none of these companies are held responsible for missing their actual SLAs, as opposed to their self-reported SLA compliance. So unless regulation gets implemented that says otherwise, there's zero incentive for any company to maintain an accurate status page.
If not happy with the results switch.
Re: AWS us-east-1 outage
#376Are the actual services down, or is it just the console and/or login page? For example, the sign-up page appears to be working: https://portal.aws.amazon.com/billing/signup#/start Are websites that run on AWS us-east up? Are the AWS CLIs working?
One of my sites went offline an hour ago because the web server stopped responding. I can't SSH into it or get any type of response. The database server in the same region and zone is continuing to run fine though.
Re: AWS us-east-1 outage
#377Earlier quoted context omitted.
I don't think engineers can believe in no-blame analysis if they know it'll harm career growth. I can't unilaterally promote John Doe, I have to convince other leaders that John would do well the next level up. And in those discussions, they could bring up "but John has caused 3 incidents this year", and honestly, maybe they'd be right.
Would they? Having 3 outages in a year sounds like an organization problem. Not enough safeguards to prevent very routine human errors. But instead of worrying about that we just assign a guy to take the fall
Re: AWS us-east-1 outage
#378Earlier quoted context omitted.
I don't think engineers can believe in no-blame analysis if they know it'll harm career growth. I can't unilaterally promote John Doe, I have to convince other leaders that John would do well the next level up. And in those discussions, they could bring up "but John has caused 3 incidents this year", and honestly, maybe they'd be right.
Would they? Having 3 outages in a year sounds like an organization problem. Not enough safeguards to prevent very routine human errors. But instead of worrying about that we just assign a guy to take the fall
Engineers should be experts and you should be able to trust them to make reasonable choices about the management of their projects.
That doesn't mean there can't be some checks in place, and it doesn't mean that all engineers should be perfect.
But you also have to acknowledge that adding all of those safeties has a cost. You can be a competent person who requires fewer safeties or less competent with more safeties.
Which one provides more value to an organization?
Re: AWS us-east-1 outage
#379Friends tell friends to pick us-east-2. Virginia is for lovers, Ohio is for availability.
If you're not multi-cloud in 2021 and are expecting 5-9's, I feel bad for you.
Re: AWS us-east-1 outage
#380Earlier quoted context omitted.
EC2 or S3 showing red in any region literally requires personal approval of the CEO of AWS.
Is this true or a joke? This sort of policy is how you destroy trust.