Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

191–200 of 1001 posts

Re: AWS us-east-1 outage

#192
post #70

This got me thinking, are there any major chat services that would go down if a particular AWS/GCP/etc data centre went down? You don't want your service to go down, plus your team's comms at the same time.

Especially if enough Amazon internal tools rely on it - would be funny if there were a repeat of the FB debacle where Amazon employees somehow couldn't communicate/get back into their offices because of the problem they were trying to fix

Last I knew, Amazon used all Microsoft stuff for business communication.

Re: AWS us-east-1 outage

#193

I worked at a company that hired an ex-Amazon engineer to work on some cloud projects. Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault. He later explained that in his…

This gets posted every time there's an AWS outage. It mind as well be a copy pasta at this point.

I mean it's true at every company I've ever worked at too. If you can lawyer incidents into not being an outage you avoid like 15 meetings with the business stakeholders about all the things we "have to do" to prevent things like this in the future that get canceled the moment they realize that how much dev/infra time it will take to implement.

Re: AWS us-east-1 outage

#195
post #113

Friends tell friends to pick us-east-2. Virginia is for lovers, Ohio is for availability.

If you're not multi-cloud in 2021 and are expecting 5-9's, I feel bad for you.

How do you become multi-cloud if your root domain is in Route53? Have Backup domains on the client side?

Re: AWS us-east-1 outage

#196
post #164

Earlier quoted context omitted.

Because Amazon has $$$$$ in their SLOs, and it costs them through the nose every minute they're down in payments made to customers and fees refunded. I trust them and most companies not to be outright fraudulent (although I'm sure some are), but it's totally understandable they'd be reticent to push the "Downtime Alert/Cost Us a Ton of Money" button until they're sure something serious is happening.

This is an incentive to dishonesty, leading to fraudulent payments and false advertising of uptime to potential customers. Hopefully it results in a class action lawsuit for enough money that Amazon decides that an automated system is better than trying to supply human judgement.

Can someone just have a site ping all the GET endpoints on the AWS API? That is very far from "automating [their entire] system" but it's better than what they're doing.

Re: AWS us-east-1 outage

#197
post #122

Earlier quoted context omitted.

Sure, but... that just raises more questions :) Taken literally what you are saying is the service could be down and an executive could override that, preventing them for paying customers for a service outage, even if the service did have an outage and the customer could prove it (screenshots, metrics from other cloud providers, many different folks see it). I'm sure there is some subtlety to this, but it does mean t…

I have no inside knowledge or anything but it seems like there are a lot of scenarios with degraded performance where people could argue about whether it really constitutes an outage.

Yep. I was an SRE who worked at Google and also launched a product on Google Cloud. We had these arguments all the time, and the contract language often provides a way for the provider to weasel out.

Re: AWS us-east-1 outage

#198

Earlier quoted context omitted.

This gets posted every time there's an AWS outage. It mind as well be a copy pasta at this point.

well, this is the first time I've seen it, so I am glad it was posted this time.

First time I've seen it too. Definitely not my first "AWS us-east-1 is down but the status board is green" thread, either.
Post reply on HN