Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

841–850 of 1001 posts

Re: AWS us-east-1 outage

#841
post #375

Earlier quoted context omitted.

How did you find a way to bring regulations into this? There are monitoring services you can pay for to keep an eye on your SLAs and your vendors'. If not happy with the results switch.

Technically, there are already regulations. SLA lies are fraud. But I'm leery of any business who's so dishonest they fear any outside oversight that brings repercussions for said dishonesty. "If not happy, switch" is silly - it's not the customer's problem. And if you're a large customer and have invested heavily in getting staff trained on AWS, you can't just move.

A) don't build a business that relies solely on existence of another B) switch to another vendor if not happy with current vendor

Really not that complicated.

Re: AWS us-east-1 outage

#842
post #788

Earlier quoted context omitted.

> This seems like an insane stance to have, it's like saying businesses should ship their own stock, using their own drivers, and their in-house made cars and planes and in-house trained pilots. > Heck, why stop at having servers on-site? Cast your own silicon waffers, after all you don't want spectrum exploits. That's an overblown argument. Nobody is saying that, but it's clear that businesses that maintain their ow…

> it's clear that businesses that maintain their own infrastructure would've avoided today's AWS' outage. Sure, that's trivially obvious. But how many other outages would they have had instead because they aren't as experienced at running this sort of infrastructure as AWS is? You seem to be arguing from the a priori assumption that rolling your own is inherently more stable than renting infra from AWS, without actua…

You and GP are making the same assumption that my DevOps engineers _aren't_ as experienced as AWS' are. There are plenty of engineers capable of maintaining an in-house infrastructure running X 9s because, again, the complexity comes from the scale AWS operates at. So we're both arguing with an a priori assumption that the grass is greener on our side.

To be fair, I'm not saying never use cloud providers. If your systems require the complexity cloud providers simplify, and you operate at a scale where it would be prohibitively expensive to maintain yourself, by all means go with a cloud provider. But it's clear that not many companies are prepared for this type of failure, and protecting against it is not trivial to accomplish. Not to mention the conceptual overhead and knowledge required with dealing with the provider's specific products, APIs, etc. Whereas maintaining these systems yourself is transferrable across any datacenter.

Re: AWS us-east-1 outage

#843

Earlier quoted context omitted.

It's popular to upvote this during outages, because it fits a narrative. The truth (as always) is more complex: * No, this isn't the broad culture. It's not even a blip. These are EXCEPTIONAL circumstances by extremely bad teams that - if and when found out - would be intervened dramatically. * The broad culture is blameless post-mortems. Not whose fault is it. But what was the problem and how to fix it. And one of t…

>* Every AWS customer has a PERSONAL health dashboard in the console that should indicate their experience. You mean the one that is down right now?

Seems like it's doing an exemplary job of indicating their experience, then.

Re: AWS us-east-1 outage

#844

After over 45 minutes https://status.aws.amazon.com/ now shows "AWS Management Console - Increased Error Rates" I guess 100% is technically an increase.

I can't remember seeing problems be more strongly worded than "Increased Error Rates" or "high error rates with S3 in us-east-1" during the infamous S3 outage of 2017 - and that was after they struggled to even update their own status page because of S3 being down. :)

During the Facebook outage FB wrote something along the lines of "We noticed that some users are experiencing issues with our apps" eventhough nothing worked anymore

Re: AWS us-east-1 outage

#845
post #217

Earlier quoted context omitted.

That sounds like the exact opposite of human-factors engineering. No one likes taking blame. But when things go sideways, people are extra spicy and defensive, which makes them clam up and often withhold useful information, which can extend the outage. No-blame analysis is a much better pattern. Everyone wins. It's about building the system that builds the system. Stuff broke; fix the stuff that broke, then fix the t…

I worked at Walmart Technology. I bravely wrote post mortem documents owning the fault of my team (100+ people), owning both technically and also culturally as their leader. I put together a plan to fix it and executed it. Thought that was the right thing to do. This happend two times in my 10 year career there. Both times I was called out as a failure in my performance eval. Second time, I resigned and told them to…

Stories like this are why I'm really glad I stopped talking to that Walmart Technology recruiter a few years ago. I love working for places where senior leadership constantly repeat war stories about "that time I broke the flagship product" to reinforce the importance of blameless postmortems. You can't fix the process if the people who report to you feel the need to lie about why things go wrong.

Re: AWS us-east-1 outage

#846

Earlier quoted context omitted.

When I was at Google I didn't have a lot of exposure to the public infra side. However I do remember back in 2008 when a colleague was working on routing side of YouTube, he made a change that cost millions of dollars in mere hours before noticing and reverting it. He mentioned this to the larger team which gave applause during a tech talk. I cannot possibly generalize the culture differences between Amazon and Googl…

While I support that, how are the people involved evaluated?

If an engineer causes an outage by mistake and then ensures that would never happen again, he has made a positive impact.

Re: AWS us-east-1 outage

#847
post #789

Earlier quoted context omitted.

Every damn Well-Architected Framework includes multi-AZ if not multi-region redundancy, and yet the single access point for their millions of customers is single-region. Facepalm in the form of $100Ms in service credits.

> Facepalm in the form of $100Ms in service credits. It was also greatly affecting Amazon.com itself. I kept getting sporadic 404 pages and one was during a purchase. Purchase history wasn't showing the product as purchased and I didn't receive an email, so I repurchased. Still no email, but the purchase didn't end in a 404, but the product still didn't show up in my purchase history. I have no idea if I purchased an…

Oh no... I think you may be in for a rough time, because I purchased something this morning and it only popped up in my orders list a few minutes ago.

Re: AWS us-east-1 outage

#848

Earlier quoted context omitted.

Maybe has something to do with CloudFront mandating certs to be in us-east-1?

YES! Why do they do that? It's so weird. I will deploy a whole config into us-west-1 or something; but then I need to create a new cert in us-east-1 JUST to let cloudfront answer an HTTPS call. So frustrating.

Agreed - in my line of work regulators want everything in the country we operate from but of course CloudFront has to be different.

Re: AWS us-east-1 outage

#849
post #245

Looks like they've acknowledged it on the status page now. https://status.aws.amazon.com/ > 8:22 AM PST We are investigating increased error rates for the AWS Management Console. > 8:26 AM PST We are experiencing API and console issues in the US-EAST-1 Region. We have identified root cause and we are actively working towards recovery. This issue is affecting the global console landing page, which is also hosted in US…

> This issue is affecting the global console landing page, which is also hosted in US-EAST-1 Even this little tidbit is a bit of a wtf for me. Why do they consider it ok to have anything hosted in a single region? At a different (unnamed) FAANG, we considered it unacceptable to have anything depend on a single region. Even the dinky little volunteer-run thing which ran https://internal.site.example/~someEngineer was…

Forget the number of regions. Monitoring for X shouldn't even be hosted on X at all...

Re: AWS us-east-1 outage

#850
I don't think AWS knows what's going on judging by their updates, yes DynamoDB might be having issues, but so is IAM it seems, we're getting issues terminating resources for example.
Post reply on HN