Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

871–880 of 1001 posts

Re: AWS us-east-1 outage

#871

Earlier quoted context omitted.

Any yet oddly enough the Earth continues to spin and the internet continues to work. I think the system we have now is necessarily the system that must exist ( in this particular case, not in all cases ). Something more centralized is destined to fail. And, while the open source nature of software introduces vulnerabilities it also fixes them.

> And, while the open source nature of software introduces vulnerabilities it also fixes them. dat gap tho... which was my point. smart black hats will be exploiting this gap, at scale. and the strategy will work because the majority of folks seem to be either lazy, ignorant or simply hurried for time. and btw your 1st sentence was rude. constructive feedback for the future

For my vote, I don't think it was rude, I think it was making a point.

Re: AWS us-east-1 outage

#872
post #734

The fun thing about these types of outages are seeing all of the people that depend upon these services with no graceful fallback. My roomba app will not even launch because of the AWS outage. I understand that the app gets "updates" from the cloud. In this case "updates" is usually promotional crap, but whatevs. However, for this to prevent the app launching in a manner that I can control my local device is total BS…

Now think of how many assets of various governments' militaries are discreetly employed as normal operational staff by FAAMG in the USA and have access to cause such events from scratch. I would imagine that the US IC (CIA/NSA) already does some free consulting for these giant companies to this end, because they are invested in that Not Being Possible (indeed, it's their job). There is a societal resilience benefit t…

> I would imagine that the US IC (CIA/NSA) already does some free consulting for these giant companies

This comment is how I know you don't work in the public sector. Those agencies' infrastructures are essentially run by contractors with a few GS personnel making bad decisions every chance they get and a few DoD personnel acting like their rank can fix technical problems.

Re: AWS us-east-1 outage

#873
post #863
post #839

Earlier quoted context omitted.

>Facepalm in the form of $100Ms in service credits. Part of me wonders how much they're actually going to pay out, given that their own status page has only indicated five services with moderate ("Increased API Error Rates") disruptions in service.

Utter lies on that page. Multiple services listed as green aren't working for me or my team.

This point is repeated often, and the incentives for Amazon to downplay the actual downtime are definitely there.

Wouldn't affected companies be incentivized to make a lawsuit about AMZ lying about status? It would be easy to prove and costly to defend from AWS standpoint.

Re: AWS us-east-1 outage

#874
post #245

Looks like they've acknowledged it on the status page now. https://status.aws.amazon.com/ > 8:22 AM PST We are investigating increased error rates for the AWS Management Console. > 8:26 AM PST We are experiencing API and console issues in the US-EAST-1 Region. We have identified root cause and we are actively working towards recovery. This issue is affecting the global console landing page, which is also hosted in US…

I like how 6 hours in: "Many services have already recovered".

Re: AWS us-east-1 outage

#875

Earlier quoted context omitted.

I firmly believe in the dictum "if you ship it you own it". That means you own all outages. It's not just an operator flubbing a command, or a bit of code that passed review when it shouldn't. It's all your dependencies that make your service work. You own ALL of them. People spend all this time threat modelling their stuff against malefactors, and yet so often people don't spend any time thinking about the threat mo…

That's a great philosophy. Ok, let's take an organization, let's call them, say Ammizzun. Totally not Amazon. Let's say you have a very aggressive hire/fire policy which worked really well in rapid scaling and growth of your company. Now you have a million odd customers highly dependent on systems that were built by people that are now one? two? three? four? hire/fire generations up-or-out or cashed-out cycles ago. S…

I worked on a project like this in government for my first job. I was the third butt in that seat in a year. Everyone associated with project that I knew there was gone by one year from my own departure date.

They are now on the 6th butt in that seat in 4 years. That poor fellow is entirely blameless for the mess that accumulated over time.

Re: AWS us-east-1 outage

#876
post #666

Earlier quoted context omitted.

>The fun thing about these types of outages are seeing all of the people that depend upon these services with no graceful fallback. Whats a graceful fallback? Switching to another hosting service when AWS goes down? Wouldn't that present another set of complications for a very small edge case at huge cost?

Usually this refers to falling back to a different region in AWS. It's typical for systems to be deployed in multiple regions due to latency concerns, but it's also important for resiliency. What you call "a very small edge case" is occurring as we speak, and if you're vulnerable to it you could be losing millions of dollars.

I heard from someone at that they couldn't switch to their fallback region because they couldn't update DNS on Route53. All the console commands and web interface were failing.

Re: AWS us-east-1 outage

#877
Anyone else having trouble with their Alexa devices? Mine are acting really wonky, couldn’t listen to NPR in the morning really messed up my routine.

Re: AWS us-east-1 outage

#878
post #877

Anyone else having trouble with their Alexa devices? Mine are acting really wonky, couldn’t listen to NPR in the morning really messed up my routine.

My Alexa-controlled lights didn’t turn on this evening.

Re: AWS us-east-1 outage

#880

I worked at a company that hired an ex-Amazon engineer to work on some cloud projects. Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault. He later explained that in his…

Former AWSser. I can totally believe that happened and continues to happen in some teams. Officially, it's not supposed to be done that way. Some AWS managers and engineers bring their corporate cultural baggage with them when they join AWS and it takes a few years to unlearn it.

Thanks for the perspective. I was beginning to regret posting this after so many people claiming this wouldn’t happen at AWS.

Amazon is a huge company so I have no doubt YMMV depending on your manager.

Post reply on HN