Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

731–740 of 1001 posts

Re: AWS us-east-1 outage

#731

Earlier quoted context omitted.

Usually this refers to falling back to a different region in AWS. It's typical for systems to be deployed in multiple regions due to latency concerns, but it's also important for resiliency. What you call "a very small edge case" is occurring as we speak, and if you're vulnerable to it you could be losing millions of dollars.

AWS itself has a huge single point of failure on us-east-1 region. Usually, if us-east-1 goes down, others soon follow. At that point, it doesn't matter how many regions you're deploying to.

My workloads on Sydney and London are unaffected. I can't speak for anywhere else.

Re: AWS us-east-1 outage

#732
post #245

Looks like they've acknowledged it on the status page now. https://status.aws.amazon.com/ > 8:22 AM PST We are investigating increased error rates for the AWS Management Console. > 8:26 AM PST We are experiencing API and console issues in the US-EAST-1 Region. We have identified root cause and we are actively working towards recovery. This issue is affecting the global console landing page, which is also hosted in US…

Ok, we've changed the URL to that from https://us-east-1.console.aws.amazon.com/console/home since the latter is still not responding.

There are also various media articles but I can't tell which ones have significant new information beyond "outage".

Re: AWS us-east-1 outage

#733

Earlier quoted context omitted.

Usually this refers to falling back to a different region in AWS. It's typical for systems to be deployed in multiple regions due to latency concerns, but it's also important for resiliency. What you call "a very small edge case" is occurring as we speak, and if you're vulnerable to it you could be losing millions of dollars.

AWS itself has a huge single point of failure on us-east-1 region. Usually, if us-east-1 goes down, others soon follow. At that point, it doesn't matter how many regions you're deploying to.

"Usually"? When has that ever happened?

Re: AWS us-east-1 outage

#734

The fun thing about these types of outages are seeing all of the people that depend upon these services with no graceful fallback. My roomba app will not even launch because of the AWS outage. I understand that the app gets "updates" from the cloud. In this case "updates" is usually promotional crap, but whatevs. However, for this to prevent the app launching in a manner that I can control my local device is total BS…

Now think of how many assets of various governments' militaries are discreetly employed as normal operational staff by FAAMG in the USA and have access to cause such events from scratch. I would imagine that the US IC (CIA/NSA) already does some free consulting for these giant companies to this end, because they are invested in that Not Being Possible (indeed, it's their job).

There is a societal resilience benefit to not having unnecessary cloud dependencies beyond the privacy stuff. It makes your society and economy more robust if you can continue in the face of remote failures/errors.

It is December 7th, after all.

Re: AWS us-east-1 outage

#735
post #328

I worked at a company that hired an ex-Amazon engineer to work on some cloud projects. Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault. He later explained that in his…

I've worked for Amazon for 4 years, including stints at AWS, and even in my current role my team is involved in LSE's. I've never seen this behavior, the general culture has been find the problem, fix it, and then do root cause analysis to avoid it again. Jeff himself has said many times in All Hands and in public "Amazon is the best place to fail". Mainly because things will break, it's not that they break that's in…

LSE?

Re: AWS us-east-1 outage

#736
post #735
post #328

Earlier quoted context omitted.

I've worked for Amazon for 4 years, including stints at AWS, and even in my current role my team is involved in LSE's. I've never seen this behavior, the general culture has been find the problem, fix it, and then do root cause analysis to avoid it again. Jeff himself has said many times in All Hands and in public "Amazon is the best place to fail". Mainly because things will break, it's not that they break that's in…

LSE?

Large Scale Event

Re: AWS us-east-1 outage

#737

Haha my developer called me in panic telling that he crashed Amazon - was doing some load tests with Lambda

Postmortem: unbounded auto-scaling of lambda combined with oversight on internal rate limits caused unforseen internal ddos.

Re: AWS us-east-1 outage

#739
post #245

Looks like they've acknowledged it on the status page now. https://status.aws.amazon.com/ > 8:22 AM PST We are investigating increased error rates for the AWS Management Console. > 8:26 AM PST We are experiencing API and console issues in the US-EAST-1 Region. We have identified root cause and we are actively working towards recovery. This issue is affecting the global console landing page, which is also hosted in US…

They are still lying about it, the issues are not only affecting the console but also AWS operations such as S3 puts. S3 still shows green.

Yep, I am seeing failures on IAM as well:

   aws iam list-policies
  
  An error occurred (503) when calling the ListPolicies operation (reached max retries: 2): Service Unavailable

Re: AWS us-east-1 outage

#740

Earlier quoted context omitted.

That's shockingly stupid. I also worked for a major Walmart IT services vendor in another life, and we always had to be careful about how we handled them, because they didn't always show a lot of respect for vendors. On another note, thanks for building some awesome stuff -- walmart.com is awesome. I have both Prime and whatever-they're-currently-calling Walmart's version and I love that Walmart doesn't appear to mix…

walmart.com user design sucks. My particular grudge right now is - I'm shopping to go pickup some stuff (and indicate "in store pickup) and each time I search for the next item, it resets that filter making me click on that filter for each item on my list.

Almost every physical-store-chain company's website makes it way too hard to do the thing I nearly always want out of their interface, which is to search the inventory of the X nearest locations. They all want to push online orders or 3rd-party-seller crap, it seems.
Post reply on HN