Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

601–610 of 1001 posts

Re: AWS us-east-1 outage

#601

https://status.aws.amazon.com hasn't been updated to reflect outage yest

There’s probably a lambda somewhere supposed to update it that is screaming into the darkness at the moment.

According to an internal message I saw, their monitoring stuff is fucked too.

Re: AWS us-east-1 outage

#603

Earlier quoted context omitted.

I firmly believe in the dictum "if you ship it you own it". That means you own all outages. It's not just an operator flubbing a command, or a bit of code that passed review when it shouldn't. It's all your dependencies that make your service work. You own ALL of them. People spend all this time threat modelling their stuff against malefactors, and yet so often people don't spend any time thinking about the threat mo…

if I were a black hat I would absolutely love GitHub and all the various language-specific package systems out there. giving me sooooo many ways to sneak arbitrary tailored malicious code into millions of installs around the world 24x7. sure, some of my attempts might get caught, or not but not lead to a valuable outcome for me. but that percentage that does? can make it worth it. its about scale and a massive parall…

Any yet oddly enough the Earth continues to spin and the internet continues to work. I think the system we have now is necessarily the system that must exist ( in this particular case, not in all cases ). Something more centralized is destined to fail. And, while the open source nature of software introduces vulnerabilities it also fixes them.

Re: AWS us-east-1 outage

#604

Earlier quoted context omitted.

I firmly believe in the dictum "if you ship it you own it". That means you own all outages. It's not just an operator flubbing a command, or a bit of code that passed review when it shouldn't. It's all your dependencies that make your service work. You own ALL of them. People spend all this time threat modelling their stuff against malefactors, and yet so often people don't spend any time thinking about the threat mo…

That's a great philosophy. Ok, let's take an organization, let's call them, say Ammizzun. Totally not Amazon. Let's say you have a very aggressive hire/fire policy which worked really well in rapid scaling and growth of your company. Now you have a million odd customers highly dependent on systems that were built by people that are now one? two? three? four? hire/fire generations up-or-out or cashed-out cycles ago. S…

A lot can go wrong as an organization grows, including loss of knowledge. At amazon "Ownership" officially rests with the non-technical money that owns voting shares. They control the board who controls the CEO. "Ownership" can be perverted to mean that you, a wage slave, are responsible for the mess that previous ICs left behind. The obvious thing to do in such a circumstance is quit (or don't apply). It is unfair and unpleasant to be treated in a way that gives you responsibility but no authority, and to participant in maintaining (and extending) that moral hazard, and as long as there are better companies you're better off working for them.

Re: AWS us-east-1 outage

#607

I worked at a company that hired an ex-Amazon engineer to work on some cloud projects. Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault. He later explained that in his…

It's popular to upvote this during outages, because it fits a narrative. The truth (as always) is more complex: * No, this isn't the broad culture. It's not even a blip. These are EXCEPTIONAL circumstances by extremely bad teams that - if and when found out - would be intervened dramatically. * The broad culture is blameless post-mortems. Not whose fault is it. But what was the problem and how to fix it. And one of t…

Hiding behind a throw away account does not help your point.

Re: AWS us-east-1 outage

#608
post #571

"some customers may experience a slight elevation in error rates" --> everything is on fire

I'm also experiencing a slight elevation in billing rates - got alarms for 10x consumption and I can't check on them... Edit: also API access is failing, terraform can't take anything down because "connection was forcibly closed"

Imagine triggering a big instance for machine learning or a huge EMR cluster that would otherwise be short lived and not being able to scale it down.

I am quite sure the AWS support will be getting many refund requests over the course of the week.

Re: AWS us-east-1 outage

#609
post #605
post #339

...and imdb

And Netflix.

And Venmo, and McDonald's, and....

This one is pretty epic (pun intended). Bad enough that Down Detector [0] shows "Reports indicate there may be a widespread outage at Amazon Web Services, which may be impacting your service." in a red alert bar at the top.

[0] - https://downdetector.com/

Post reply on HN