https://status.aws.amazon.com hasn't been updated to reflect outage yest
According to an internal message I saw, their monitoring stuff is fucked too.
601–610 of 1001 posts
https://status.aws.amazon.com hasn't been updated to reflect outage yest
According to an internal message I saw, their monitoring stuff is fucked too.
Earlier quoted context omitted.
I firmly believe in the dictum "if you ship it you own it". That means you own all outages. It's not just an operator flubbing a command, or a bit of code that passed review when it shouldn't. It's all your dependencies that make your service work. You own ALL of them. People spend all this time threat modelling their stuff against malefactors, and yet so often people don't spend any time thinking about the threat mo…
if I were a black hat I would absolutely love GitHub and all the various language-specific package systems out there. giving me sooooo many ways to sneak arbitrary tailored malicious code into millions of installs around the world 24x7. sure, some of my attempts might get caught, or not but not lead to a valuable outcome for me. but that percentage that does? can make it worth it. its about scale and a massive parall…
Earlier quoted context omitted.
I firmly believe in the dictum "if you ship it you own it". That means you own all outages. It's not just an operator flubbing a command, or a bit of code that passed review when it shouldn't. It's all your dependencies that make your service work. You own ALL of them. People spend all this time threat modelling their stuff against malefactors, and yet so often people don't spend any time thinking about the threat mo…
That's a great philosophy. Ok, let's take an organization, let's call them, say Ammizzun. Totally not Amazon. Let's say you have a very aggressive hire/fire policy which worked really well in rapid scaling and growth of your company. Now you have a million odd customers highly dependent on systems that were built by people that are now one? two? three? four? hire/fire generations up-or-out or cashed-out cycles ago. S…
...and imdb
I worked at a company that hired an ex-Amazon engineer to work on some cloud projects. Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault. He later explained that in his…
It's popular to upvote this during outages, because it fits a narrative. The truth (as always) is more complex: * No, this isn't the broad culture. It's not even a blip. These are EXCEPTIONAL circumstances by extremely bad teams that - if and when found out - would be intervened dramatically. * The broad culture is blameless post-mortems. Not whose fault is it. But what was the problem and how to fix it. And one of t…
"some customers may experience a slight elevation in error rates" --> everything is on fire
I'm also experiencing a slight elevation in billing rates - got alarms for 10x consumption and I can't check on them... Edit: also API access is failing, terraform can't take anything down because "connection was forcibly closed"
I am quite sure the AWS support will be getting many refund requests over the course of the week.
...and imdb
And Netflix.
This one is pretty epic (pun intended). Bad enough that Down Detector [0] shows "Reports indicate there may be a widespread outage at Amazon Web Services, which may be impacting your service." in a red alert bar at the top.
"some customers may experience a slight elevation in error rates" --> everything is on fire