Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

231–240 of 1001 posts

Re: AWS us-east-1 outage

#232
post #38
post #13

I love that every time this happens, 100% of the services on https://status.aws.amazon.com are green.

Those five 9s don't come easy. Sometimes you have to prop them up :)

Every time someone asks to update the status page, managers say "nein"

Re: AWS us-east-1 outage

#234
post #195

Earlier quoted context omitted.

If you're not multi-cloud in 2021 and are expecting 5-9's, I feel bad for you.

How do you become multi-cloud if your root domain is in Route53? Have Backup domains on the client side?

dns records should be synced to secondary provider, and that provider added to your secondary/tertiary domain dns.

Multi-provider dns is a solved problem.

Re: AWS us-east-1 outage

#235
post #151

Earlier quoted context omitted.

Because Amazon has $$$$$ in their SLOs, and it costs them through the nose every minute they're down in payments made to customers and fees refunded. I trust them and most companies not to be outright fraudulent (although I'm sure some are), but it's totally understandable they'd be reticent to push the "Downtime Alert/Cost Us a Ton of Money" button until they're sure something serious is happening.

It should be costing them trust not to push it when they should though. A trustworthy company will err on the side of pushing it. AWS is a near-monopoly, so their unprofessional business practices have still yet to cost them.

> It should be costing them trust not to push it when they should though.

This is what Amazon, the startup, understood.

Step 1: Always make it right and make the customer happy, even if it hurts in $.

Step 2: If you find you're losing too much money over a particular issue, fix the issue.

Amazon, one of the world's largest companies, seems to have forgotten that the risk of not reporting accurately isn't money, but breaking the feedback chain. Once you start gaming metrics, no leaders know what's really important to work on internally, because no leaders know what the actual issues are. It's late Soviet Union in a nutshell. If everyone is gaming the system at all levels, then eventually the ability to objectively execute decreases, because effort is misallocated due to misunderstanding.

Re: AWS us-east-1 outage

#236
post #122

Earlier quoted context omitted.

Sure, but... that just raises more questions :) Taken literally what you are saying is the service could be down and an executive could override that, preventing them for paying customers for a service outage, even if the service did have an outage and the customer could prove it (screenshots, metrics from other cloud providers, many different folks see it). I'm sure there is some subtlety to this, but it does mean t…

I have no inside knowledge or anything but it seems like there are a lot of scenarios with degraded performance where people could argue about whether it really constitutes an outage.

One time gcp argued that since they did return 404s on gcs for a few hours that wasn’t an uptime/latency sla violation so we were not entitled to refund (tho they refunded us anyway)

Re: AWS us-east-1 outage

#237
post #13

I love that every time this happens, 100% of the services on https://status.aws.amazon.com are green.

> Goodhart's Law is expressed simply as: “When a measure becomes a target, it ceases to be a good measure.” It’s very frustrating. Why even have them?

Because "uptime" and "nines" became a marketing term. Simple as that. But the problem is that any public-facing measure of availability becomes a defacto marketing term.

Re: AWS us-east-1 outage

#239

I worked at a company that hired an ex-Amazon engineer to work on some cloud projects. Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault. He later explained that in his…

Every incident review meeting I've ever been in starts out like, _"This meeting isn't to place blame..."_, then, 5 minutes later, it turns into the Blame Game.

Re: AWS us-east-1 outage

#240
post #146

Earlier quoted context omitted.

It's not really dishonest though because there is nuance. Most everything in EC2 is still working it seems, just the console is down. So is it really down? It should probably be yellow but not red.

if you cannot access the control plane to create or destroy resources, it is down (partial availability). The jobs that are running are basically zombies.

Depending the workload being run users may or may not notice. Should be Yellow at a minimum.
Post reply on HN