Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

161–170 of 1001 posts

Re: AWS us-east-1 outage

#161

Earlier quoted context omitted.

Being dishonest about SLAs seems to bear zero cost in this case?

It's not really dishonest though because there is nuance. Most everything in EC2 is still working it seems, just the console is down. So is it really down? It should probably be yellow but not red.

"Good at finding excuses" is not the same thing as "honest."

Re: AWS us-east-1 outage

#162
post #13

I love that every time this happens, 100% of the services on https://status.aws.amazon.com are green.

Makes you wonder if they have to manually update the page when outages occur. That'd be a pretty bad way to go, so I'd hope not. Maybe the code to automatically update the page is in us-east-1? :)

Word on the street is the status page is just a JPG

Re: AWS us-east-1 outage

#164

Earlier quoted context omitted.

I don't see why they couldn't provide an error rate graph like Reddit[0] or simply make services yellow saying "increased error rate detected, investigating..." 0: https://www.redditstatus.com/#system-metrics

Because Amazon has $$$$$ in their SLOs, and it costs them through the nose every minute they're down in payments made to customers and fees refunded. I trust them and most companies not to be outright fraudulent (although I'm sure some are), but it's totally understandable they'd be reticent to push the "Downtime Alert/Cost Us a Ton of Money" button until they're sure something serious is happening.

This is an incentive to dishonesty, leading to fraudulent payments and false advertising of uptime to potential customers.

Hopefully it results in a class action lawsuit for enough money that Amazon decides that an automated system is better than trying to supply human judgement.

Re: AWS us-east-1 outage

#165
post #122

Earlier quoted context omitted.

From what I've heard it's mostly true. Not only the CEO but a few SVPs can approve it, but yes a human must approve the update and it must be a high level exec. Part of the reason is because their SLAs are based on that dashboard, and that dashboard going red has a financial cost to AWS, so like any financial cost, it needs approval.

Sure, but... that just raises more questions :) Taken literally what you are saying is the service could be down and an executive could override that, preventing them for paying customers for a service outage, even if the service did have an outage and the customer could prove it (screenshots, metrics from other cloud providers, many different folks see it). I'm sure there is some subtlety to this, but it does mean t…

I have no inside knowledge or anything but it seems like there are a lot of scenarios with degraded performance where people could argue about whether it really constitutes an outage.

Re: AWS us-east-1 outage

#166

I worked at a company that hired an ex-Amazon engineer to work on some cloud projects. Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault. He later explained that in his…

Sometimes, these large companies tack on too much "necessary" incident "remediation" actions with Arbitrary Due Date SLAs that completely wrench any ongoing work. And ongoing, strategically defined ""muh high impact"" projects are what get you promoted, not doing incident remediations.

When you get to the level you want, you get to not really give a shit and actually do The Right Thing. However, for all of the engineers clamoring to get out of the intermediate brick laying trenches, opening an incident can have pervasive incentives.

Re: AWS us-east-1 outage

#170

This got me thinking, are there any major chat services that would go down if a particular AWS/GCP/etc data centre went down? You don't want your service to go down, plus your team's comms at the same time.

Slack is pretty much full AWS, I've been switched over to Teams so I can't check.

Slack is working fine for me
Post reply on HN