Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

261–270 of 1001 posts

Re: AWS us-east-1 outage

#261

Have folks considered a class-action lawsuit against these blatantly fraudulent SLAs to recoup costs?

In my experience, despite whatever is published, companies will private acknowledge and pay their SLA terms. (Which still only gets you, like, one day's worth of reimbursement if you're lucky.)

Re: AWS us-east-1 outage

#262
post #62

Earlier quoted context omitted.

They can do this without an alliance. They very intentionally choose not to do it. Every major company has moved away from having accurate status pages.

It's because none of these companies are held responsible for missing their actual SLAs, as opposed to their self-reported SLA compliance. So unless regulation gets implemented that says otherwise, there's zero incentive for any company to maintain an accurate status page.

The only uptimes I'm concerned with are my own services that my own monitoring keeps on top of, this varies - if the monitoring page goes down for 10 seconds I'm not worried, if one leg of a smpte-2022-7 is down for a second that's fine, if it keeps going down for a second that's a concern, etc.

If something I'm responsible goes down to the point that my stakeholders are complaining (which is something seriously wrong), they are not going to be happy with "oh the cloud was down, not my fault"

If AWS is down or not it meaningless to me, if my service running on AWS is down or not is the key metric.

If a service is down and I can't get into it, then chatter on things like outages mailing list, or HN, will let me know if it's yet another cloud failure, or if it's something that's affecting my machine only.

Re: AWS us-east-1 outage

#264
post #13

I love that every time this happens, 100% of the services on https://status.aws.amazon.com are green.

Are they lying, or just prioritizing their own services?

https://music.amazon.co.uk is giving me an error since about 16:30 GMT

"We are experiencing an error. Our apologies – We will be back up soon."

Re: AWS us-east-1 outage

#266
post #217

I worked at a company that hired an ex-Amazon engineer to work on some cloud projects. Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault. He later explained that in his…

That sounds like the exact opposite of human-factors engineering. No one likes taking blame. But when things go sideways, people are extra spicy and defensive, which makes them clam up and often withhold useful information, which can extend the outage. No-blame analysis is a much better pattern. Everyone wins. It's about building the system that builds the system. Stuff broke; fix the stuff that broke, then fix the t…

> No-blame analysis is a much better pattern. Everyone wins. It's about building the system that builds the system. Stuff broke; fix the stuff that broke, then fix the things that let stuff break.

Yea, except it doesn't work in practice. I work with a lot of people who come from places with "blameless" post-mortem 'culture' and they've evangelized such a thing extensively.

You know what all those people have proven themselves to really excel at? Blaming people.

Re: AWS us-east-1 outage

#267
post #13

I love that every time this happens, 100% of the services on https://status.aws.amazon.com are green.

I don't see why they couldn't provide an error rate graph like Reddit[0] or simply make services yellow saying "increased error rate detected, investigating..." 0: https://www.redditstatus.com/#system-metrics

The more transparency you give; the harder it is to control the narrative. They have a general reputation for reliability; and exposing just how many actual errors/failures there are (that generally don't effect a large swath of users/usecases) would do hurt that reputation for minimal gain.

Re: AWS us-east-1 outage

#269

I worked at a company that hired an ex-Amazon engineer to work on some cloud projects. Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault. He later explained that in his…

What if they just can't access the console to update the status page...

Re: AWS us-east-1 outage

#270

I worked at a company that hired an ex-Amazon engineer to work on some cloud projects. Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault. He later explained that in his…

This gets posted every time there's an AWS outage. It mind as well be a copy pasta at this point.

Sorry. I'm probably to blame because I've posted this a couple times on HN before.

It strikes a nerve with me because it caused so much trouble for everyone around him. He had other personal issues, though, so I should probably clarify that I'm not entirely blaming Amazon for his habits. Though his time at Amazon clearly did exacerbate his personal issues.

Post reply on HN