Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

501–510 of 1001 posts

Re: AWS us-east-1 outage

#501
post #62

Earlier quoted context omitted.

They can do this without an alliance. They very intentionally choose not to do it. Every major company has moved away from having accurate status pages.

Steam has a great status page, companies like that and cloudfare will eat Alphabet's lunch in the next 17-18 years.

That's a tight time frame a long way off. How'd you arrive at 17-18?

Re: AWS us-east-1 outage

#502

Earlier quoted context omitted.

"Fixed a bug that could cause [adverse behavior affecting 100% of the user base] for some users"

"some" as in "not all". I'm sure there are some tiny tiny sites that were unaffected because no one went to them during the outage.

"some" technically includes "all", doesn't it? It excludes "none", I suppose, but why should it exclude "all" (except if "all" equals "none")?

Re: AWS us-east-1 outage

#503
The blatant status page lies are getting absolutely ridiculous. How many hours does a service need to be totally down until it gets properly labelled as a "disruption"?

Re: AWS us-east-1 outage

#504
post #362

Earlier quoted context omitted.

Does this imply Virginia is Godless?

Virginia's actual motto is "Sic semper tyrannis". What's more tyrannical than an omnipotent being that will condemn you to eternal torment if you don't worship them and follow their laws.

I mean, most people are okay with dogs and Seattle.

Re: AWS us-east-1 outage

#505
post #13

I love that every time this happens, 100% of the services on https://status.aws.amazon.com are green.

When I worked there it required the signoff of both your VP-level executive and the comms team to update the status page. I do not believe I ever received said signoff before the issues were resolved.

Re: AWS us-east-1 outage

#506

Earlier quoted context omitted.

Cool. Now let's have a race to see who can triple their capacity the fastest. (Note: I don't use AWS, so I can actually do it)

Why would I want to triple my capacity? Most people don't need to scale to a billion users overnight.

Many B2B-type applications have a lot of usage during the workday and minimal usage outside of it. No reason to keep all that capacity running 24/7 when you only need most of it for ~8 hours per weekday. The cloud is perfect for that use case.

Re: AWS us-east-1 outage

#507

I worked at a company that hired an ex-Amazon engineer to work on some cloud projects. Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault. He later explained that in his…

It's popular to upvote this during outages, because it fits a narrative. The truth (as always) is more complex: * No, this isn't the broad culture. It's not even a blip. These are EXCEPTIONAL circumstances by extremely bad teams that - if and when found out - would be intervened dramatically. * The broad culture is blameless post-mortems. Not whose fault is it. But what was the problem and how to fix it. And one of t…

I have in the past directed users here on HN who were complaining about https://status.aws.amazon.com to the Personal Health Dashboard at https://phd.aws.amazon.com/ as well. Unfortunately even though the account I was logged into this time only has a single S3 bucket in the EU, billed through the EU and with zero direct dependancies on the US the personal health dashboard was ALSO throwing "The request processing has failed because of an unknown error" messages. Whatever the problem was this time it had global effects for the majority of users of the Console, the internet noticed for over 30 minutes before either the status page or the PHD were able to report it. There will be no explanation and the official status page logs will say there was "increased API failure rates" for an hour.

Now i guess its possible that the 1000s and 1000s of us who noticed and commented are some tiny fraction of the user base but if thats so you could at least publish a follow up like other vendors do that says something like 0.00001% of API requests failed effecting an estimated 0.001% of our users at the time.

Re: AWS us-east-1 outage

#509

I worked at a company that hired an ex-Amazon engineer to work on some cloud projects. Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault. He later explained that in his…

It's popular to upvote this during outages, because it fits a narrative. The truth (as always) is more complex: * No, this isn't the broad culture. It's not even a blip. These are EXCEPTIONAL circumstances by extremely bad teams that - if and when found out - would be intervened dramatically. * The broad culture is blameless post-mortems. Not whose fault is it. But what was the problem and how to fix it. And one of t…

Can’t comment on most of your post but I know a lot of Amazon engineers who think of the CoE process (Correction of Error, what other companies would call a postmortem) as punitive
Post reply on HN