Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

521–530 of 1001 posts

Re: AWS us-east-1 outage

#521

Earlier quoted context omitted.

Should I be penalized if an upstream dependency, owned by another team, fails? Did I lack due diligence in choosing to accept the risk that the other team couldn't deliver? These are real problems in the micro-services world, especially since I own UI and there are dozens of teams pumping out services, and I'm at the mercy of all of them. The best I can do is gracefully fail when services don't function in a healthy…

You and many others here may be conflating two concepts which are actually quite separate. Taking blame is a purely punitive action and solves nothing. Taking responsibility means it's your job to correct the problem. I find that the more "political" the culture in the organization is, the more likely everyone is to search for a scapegoat to protect their own image when a mistake happens. The higher you go up in the…

Every argument I have on the internet is between prescriptive and descriptive language.

People tend to believe that if you can describe a problem that means you can prescribe a solution. Often times, the only way to survive is to make it clear that the first thing you are doing is describing the problem.

After you do that, and it's clear that's all you are doing, then you follow up with a prescriptive description where you place clearly what could be done to manage a future scenario.

If you don't create this bright line, you create a confused interpretation.

Re: AWS us-east-1 outage

#523
post #362

Earlier quoted context omitted.

Does this imply Virginia is Godless?

Virginia's actual motto is "Sic semper tyrannis". What's more tyrannical than an omnipotent being that will condemn you to eternal torment if you don't worship them and follow their laws.

I think I should add state motto to my data center consideration matrix.

Re: AWS us-east-1 outage

#524
post #13

I love that every time this happens, 100% of the services on https://status.aws.amazon.com are green.

"some customers may experience a slight elevation in error rates" --> everything is on fire, absolutely nothing works

Re: AWS us-east-1 outage

#525

Earlier quoted context omitted.

It's popular to upvote this during outages, because it fits a narrative. The truth (as always) is more complex: * No, this isn't the broad culture. It's not even a blip. These are EXCEPTIONAL circumstances by extremely bad teams that - if and when found out - would be intervened dramatically. * The broad culture is blameless post-mortems. Not whose fault is it. But what was the problem and how to fix it. And one of t…

I haven't asked AWS employees specifically about blameless postmortems, but several of them have personally corroborated that the culture tends towards being adversarial and "performance focused." That's a tough environment for blameless debugging and postmoretems. Like if I heard that someone has a rain forest tree-frog living happily in their outdoor Arizona cactus garden, I have doubts.

When I was at Google I didn't have a lot of exposure to the public infra side. However I do remember back in 2008 when a colleague was working on routing side of YouTube, he made a change that cost millions of dollars in mere hours before noticing and reverting it. He mentioned this to the larger team which gave applause during a tech talk. I cannot possibly generalize the culture differences between Amazon and Google, but at least in that one moment, the Google culture seemed to support that errors happen, they get noticed, and fixed without harming the perceived performance of those responsible.

Re: AWS us-east-1 outage

#526
post #294

Earlier quoted context omitted.

This is the exact opposite of my experience at AWS. Amazon is all about blameless fact finding when it comes to root cause analysis. Your company just hired a not so great engineer or misunderstood him.

Adding my piece of anecdata to this.. the process is quite blameless. If a postmortem seems like it points blame, this is pointed out and removed.

Blameless, maybe, but not repercussion-less. A bad CoE was liable to upend the team's entire roadmap and put their existing goals at risk. To be fair, management was fairly receptive to "we need to throw out the roadmap and push our launch out to the following reinvent", but it wasn't an easy position for teams to be in.

Re: AWS us-east-1 outage

#528

Earlier quoted context omitted.

It's popular to upvote this during outages, because it fits a narrative. The truth (as always) is more complex: * No, this isn't the broad culture. It's not even a blip. These are EXCEPTIONAL circumstances by extremely bad teams that - if and when found out - would be intervened dramatically. * The broad culture is blameless post-mortems. Not whose fault is it. But what was the problem and how to fix it. And one of t…

Because OTHERWISE people might think AMAZON is a DYSFUNCTIONAL company that is beginning to CRATER under its HORRIBLE work culture and constant H/FIRE cycle. See, AWS is basically turning into a long standing utility that needs to be reliable. Hey, do most institutions like that completely turn over their staff every three years? Yeah, no. Great for building it out and grabbing market share. Maybe not for being the b…

> Maybe not for being the basis of a reliable substrate of the modern internet.

Maybe THEY will go to a COMPETITOR and THINGS MOVE ON if it's THAT BAD. I wasn't sure what the pattern for all caps was, so just giving it a shot there. Apologies if it's incorrect.

Re: AWS us-east-1 outage

#529
post #212

A former colleague told me years ago that us-east-1 is basically the guinea pig where changes get tested before being rolled out to the other regions, and as a result is less stable than the others. Does anyone know if there's any truth to this?

The guideline has been to deploy to it last.

If the team follows pipeline best practices, they are supposed to deploy to a single small region first, wait 24 hours, and then deploy to more, wait more, and deploy to more, until finally deploying to us-east-1.

Re: AWS us-east-1 outage

#530

No wonder I could not read books from amazon all of a sudden, what about their cloud-based redundancy design?

The book preview webservice or actual ebooks (kindle, etc)?

For me, it's been that I am unable to download books to the kindle app on my computer
Post reply on HN