Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

571–580 of 1001 posts

Re: AWS us-east-1 outage

#571

"some customers may experience a slight elevation in error rates" --> everything is on fire

I'm also experiencing a slight elevation in billing rates - got alarms for 10x consumption and I can't check on them... Edit: also API access is failing, terraform can't take anything down because "connection was forcibly closed"

Re: AWS us-east-1 outage

#572
post #62

Earlier quoted context omitted.

They can do this without an alliance. They very intentionally choose not to do it. Every major company has moved away from having accurate status pages.

It's because none of these companies are held responsible for missing their actual SLAs, as opposed to their self-reported SLA compliance. So unless regulation gets implemented that says otherwise, there's zero incentive for any company to maintain an accurate status page.

I wonder if there could be profitable play where an organization monitors SLA compliance, and then produces a batch of lawsuits or class action suit on behalf of all of its members when the SLA is violated.

Re: AWS us-east-1 outage

#573
post #164

Earlier quoted context omitted.

This is an incentive to dishonesty, leading to fraudulent payments and false advertising of uptime to potential customers. Hopefully it results in a class action lawsuit for enough money that Amazon decides that an automated system is better than trying to supply human judgement.

Can someone just have a site ping all the GET endpoints on the AWS API? That is very far from "automating [their entire] system" but it's better than what they're doing.

Something like this? https://stop.lying.cloud/

Re: AWS us-east-1 outage

#574

Earlier quoted context omitted.

It's popular to upvote this during outages, because it fits a narrative. The truth (as always) is more complex: * No, this isn't the broad culture. It's not even a blip. These are EXCEPTIONAL circumstances by extremely bad teams that - if and when found out - would be intervened dramatically. * The broad culture is blameless post-mortems. Not whose fault is it. But what was the problem and how to fix it. And one of t…

Well, the narrative is sort of what Amazon is asking for, heh? The whole us-east-1 management console is gone, what is Amazon posting for the management console on their website? "Service degradation" It's not a degradation if it's outright down. Use the red status a little bit more often, this is a "disruption", not a "degradation".

Yeah no kidding. Is there a ratio of how many people it has to be working for to be in yellow rather than red? Some internal person going “it works on my machine” while 99% of customers are down.

Re: AWS us-east-1 outage

#575
post #517

Earlier quoted context omitted.

What's a "bin" in this context?

Literally just a bin in a fulfillment warehouse. An amazon listing doesn't guarantee a particular SKU.

Ah, whew. That's what I thought. Thanks! I asked because we make warehouse and retail management systems and every vendor or customer seems to give every word their own meanings (e.g., we use "bin" in our discounts engine to be a collection of products eligible for discounts, and "barcode" has at least three meanings depending on to whom you're speaking).

Re: AWS us-east-1 outage

#576

Earlier quoted context omitted.

I've always wondered why services are not counted down more often. Is there some sliver of customers who have access to the management console for example? An increase in error rates - no biggie, any large system is going to have errors. But when 80%+ of customers loads in the region are impacted (cross availability zones for whatever good those do) - that counts as down doesn't it? Error rates in one AZ - degraded.…

yes there were. I'm from central europe and we were at least able to get some pages of the console in us-east-1 -but i assume this was more caching related. Even though the console loaded and worked for listing some entries - we weren't able to post a support case nor viewing SQS messages etc. So i aggree that degraded is not the proper wording - but it's / was not completly vanished. so.... hard to tell what is an c…

From France, when I connect to "my personal health dashboard" in eu-west-3, it says several services are having "issues" in us-east-1.

To your point, for support center (which doesn't show a region) it says:

Description

Increased Error Rates

[09:01 AM PST] We are investigating increased error rates for the Support Center console and Support API in the US-EAST-1 Region.

[09:26 AM PST] We can confirm increased error rates for the Support Center console and Support API in the US-EAST-1 Region. We have identified the root cause of the issue and are working towards resolution.

Re: AWS us-east-1 outage

#577

Earlier quoted context omitted.

I've always wondered why services are not counted down more often. Is there some sliver of customers who have access to the management console for example? An increase in error rates - no biggie, any large system is going to have errors. But when 80%+ of customers loads in the region are impacted (cross availability zones for whatever good those do) - that counts as down doesn't it? Error rates in one AZ - degraded.…

SLAs. Officially acknowledging an incident means that they now have to issue the SLA credits.

The outage dashboard is normally only updated if a certain $X percent of hosts / service is down. If the EC2 section were updated every time a rack in a datacenter went down, it would be red 24x7.

It's only updated when a large percentage of customers are impacted, and most of the time this number is less than what the HN echo chamber makes it appear to be.

Re: AWS us-east-1 outage

#579

Earlier quoted context omitted.

I don't see why they couldn't provide an error rate graph like Reddit[0] or simply make services yellow saying "increased error rate detected, investigating..." 0: https://www.redditstatus.com/#system-metrics

Because Amazon has $$$$$ in their SLOs, and it costs them through the nose every minute they're down in payments made to customers and fees refunded. I trust them and most companies not to be outright fraudulent (although I'm sure some are), but it's totally understandable they'd be reticent to push the "Downtime Alert/Cost Us a Ton of Money" button until they're sure something serious is happening.

I can Google and see how many apps, games, or other services are down. So them not "pushing some buttons" to confirm it isn't fooling anyone.

Re: AWS us-east-1 outage

#580

Funny, I just asked Alexa to set a timer and she said there was a problem doing that. Apparently timers require functioning us-east-1 now.

I can’t turn on my lights… the future is weird

And that is why my lighting automation has a baseline req that it works 100% without the internet and preferably without a central controller.
Post reply on HN