If you have trouble logging in to AWS Console, you can use a regional console endpoint such as https://eu-central-1.console.aws.amazon.com/
AWS us-east-1 outage
231–240 of 1001 posts
Re: AWS us-east-1 outage
#232Re: AWS us-east-1 outage
#233Just curious does it still make sense to claim that up time is X numbers of 9? (e.g. 99.999%)
Re: AWS us-east-1 outage
#234Earlier quoted context omitted.
If you're not multi-cloud in 2021 and are expecting 5-9's, I feel bad for you.
How do you become multi-cloud if your root domain is in Route53? Have Backup domains on the client side?
Multi-provider dns is a solved problem.
Re: AWS us-east-1 outage
#235Earlier quoted context omitted.
Because Amazon has $$$$$ in their SLOs, and it costs them through the nose every minute they're down in payments made to customers and fees refunded. I trust them and most companies not to be outright fraudulent (although I'm sure some are), but it's totally understandable they'd be reticent to push the "Downtime Alert/Cost Us a Ton of Money" button until they're sure something serious is happening.
It should be costing them trust not to push it when they should though. A trustworthy company will err on the side of pushing it. AWS is a near-monopoly, so their unprofessional business practices have still yet to cost them.
This is what Amazon, the startup, understood.
Step 1: Always make it right and make the customer happy, even if it hurts in $.
Step 2: If you find you're losing too much money over a particular issue, fix the issue.
Amazon, one of the world's largest companies, seems to have forgotten that the risk of not reporting accurately isn't money, but breaking the feedback chain. Once you start gaming metrics, no leaders know what's really important to work on internally, because no leaders know what the actual issues are. It's late Soviet Union in a nutshell. If everyone is gaming the system at all levels, then eventually the ability to objectively execute decreases, because effort is misallocated due to misunderstanding.
Re: AWS us-east-1 outage
#236Earlier quoted context omitted.
Sure, but... that just raises more questions :) Taken literally what you are saying is the service could be down and an executive could override that, preventing them for paying customers for a service outage, even if the service did have an outage and the customer could prove it (screenshots, metrics from other cloud providers, many different folks see it). I'm sure there is some subtlety to this, but it does mean t…
I have no inside knowledge or anything but it seems like there are a lot of scenarios with degraded performance where people could argue about whether it really constitutes an outage.
Re: AWS us-east-1 outage
#237I love that every time this happens, 100% of the services on https://status.aws.amazon.com are green.
> Goodhart's Law is expressed simply as: “When a measure becomes a target, it ceases to be a good measure.” It’s very frustrating. Why even have them?
Re: AWS us-east-1 outage
#238Re: AWS us-east-1 outage
#239I worked at a company that hired an ex-Amazon engineer to work on some cloud projects. Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault. He later explained that in his…
Re: AWS us-east-1 outage
#240Earlier quoted context omitted.
It's not really dishonest though because there is nuance. Most everything in EC2 is still working it seems, just the console is down. So is it really down? It should probably be yellow but not red.
if you cannot access the control plane to create or destroy resources, it is down (partial availability). The jobs that are running are basically zombies.