Live data from Hacker News

Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

aws.amazon.com

371–380 of 410 posts

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#371
post #320
post #278

Complex systems are really really hard. I'm not a big fan of seeing all these folks bash AWS for this, and not really understanding the complexity or nastiness of situations like this. Running the kind of services they do for the kind of customers, this is a VERY hard problem. We ran into a very similar issue, but at the database layer in our company literally 2 weeks ago, where connections to our MySQL exploded and…

Excuse me, do we need all that complexity? Telling that it is "hard" is justifiable? It is naive to assume people bashing AWS are uncapable to running things better, cheaper, faster, across many other vendors, on-prem, colocation or what not. > Outrage is the easy response. That is what made AWS get the marketshare it has now in the first place, the easy responses. The main selling point of AWS in the beginning was "…

> AWS started making things more complex than it should

I don’t think this is fair for a couple reasons:

1. AWS would have had to scale regardless just because of the number of customers. Even without adding features. This means many data centers, complex virtual networking, internal networks, etc. These are solving very real problems that happen when you have millions of virtual servers.

2. AWS hosts many large, complex systems like Netflix. Companies like Netflix are going to require more advanced features out of AWS, and this will result in more features being added. While this is added complexity, it’s also solving a customer problem.

My point is that complexity is inherent to the benefits of the platform.

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#372

Earlier quoted context omitted.

> I wish we would just throw up a generic "Shit's Fucked Up. We Don't Know Why Yet, But We're Working On It" message. I gotta say, the implication that you can't register an outage until you know why it happened is pretty damning. The status page is where we look to see if services are effected, if that information can't be shared there until you understand the cause, that's very broken. The AWS status page has becom…

Can you please help me understand why you, and everyone else, are so passionate about the status page? I get that it not being updated is an annoyance, but I cannot figure out why it is the single most discussed thing about this whole event. I mean, entire services were out for almost an entire day, and if you read HN threads it would seem that nobody even cares about lost revenue/productivity, downtime, etc. The vas…

https://aws.amazon.com/it/legal/service-level-agreements/

There's literally millions on the line.

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#373
post #340

Hm. This post does not seem to acknowledge what I saw. Multiple hours of rate-limiting kicking in when trying to talk to S3 (eu-west-1). After the incident everything works fine without any remediations done on our end.

eu-west-1 was not impacted by this event. I’m assuming you saw 503 Slowdown responses, which are non-exceptional and happen for a multitude of reasons.

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#374
post #340

Hm. This post does not seem to acknowledge what I saw. Multiple hours of rate-limiting kicking in when trying to talk to S3 (eu-west-1). After the incident everything works fine without any remediations done on our end.

eu-west-1 was not impacted by this event. I’m assuming you saw 503 Slowdown responses, which are non-exceptional and happen for a multitude of reasons.

I see. Then I suppose that was an unfortunately timed happenstance (and we should look into that more closely).

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#375
post #266

Earlier quoted context omitted.

> System dynamics are hard. And have to be actually tested. Most of them are designs based on nothing but uninformed intuition. There is an art to back pressure and keeping pipelines optimally utilized. Queueing doesn’t work like you think until you really know.

Why is this hard, and can’t just be written down somewhere as part of the engineering discipline? This aspect of systems in 2021 really shouldn’t be an “art.”

Because """software engineering""" is a joke.

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#376

Earlier quoted context omitted.

> I wish we would just throw up a generic "Shit's Fucked Up. We Don't Know Why Yet, But We're Working On It" message. I gotta say, the implication that you can't register an outage until you know why it happened is pretty damning. The status page is where we look to see if services are effected, if that information can't be shared there until you understand the cause, that's very broken. The AWS status page has becom…

Can you please help me understand why you, and everyone else, are so passionate about the status page? I get that it not being updated is an annoyance, but I cannot figure out why it is the single most discussed thing about this whole event. I mean, entire services were out for almost an entire day, and if you read HN threads it would seem that nobody even cares about lost revenue/productivity, downtime, etc. The vas…

> Can you please help me understand why you, and everyone else, are so passionate about the status page?

I don't think people are "passionate about status page." I think people are unhappy with someone they are supposed to trust straight up lying to their face.

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#377

I’ve been running platform teams on aws now for 10 years, and working in aws for 13. For anyone looking for guidance on how to avoid this, here’s the advice I give startups I advise. First, if you can, avoid us-east-1. Yes, you’ll miss new features, but it’s also the least stable region. Second, go multi AZ for production workloads. Safety of your customer’s data is your ethical responsibility. Protect it, back it up…

Hey Wes! I upvoted your comment before I noticed your handle. +1 insightful, as usual

Brown nose

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#379
post #278

Complex systems are really really hard. I'm not a big fan of seeing all these folks bash AWS for this, and not really understanding the complexity or nastiness of situations like this. Running the kind of services they do for the kind of customers, this is a VERY hard problem. We ran into a very similar issue, but at the database layer in our company literally 2 weeks ago, where connections to our MySQL exploded and…

But Amazon advertises that they DO understand the complexity of this, and that their understanding, knowledge and experience is so deep that they are a safe place to put your critical applications, and so you should pay them lots of money to do so.

Totally understand that complex systems behave in incomprehensible ways (hopefully only temporarily incomprehensible). But they're selling people on the idea of trading your complex system, for their far more complex system that they manage with such great expertise that it is more reliable.

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#380
post #348

Earlier quoted context omitted.

> Outrage is the easy response. Empathy and learning is the valuable one. I'm outraged that AWS, as a company policy, continues to lie about the status of their systems during outages, making it hard for me to communicate to my stakeholders. Empathy? For AWS? AWS is part a mega corporation that is closing in on 2 TRILLION dollars in market cap. It's not a person. I can empathize with individuals who work for AWS but…

It seems obvious to me that they're specifically talking about having empathy for the people who work there, the people who designed and built these systems and yes, empathy even for the people who might not be sure what to put on their absolutely humongous status page until they're sure.

But I don’t see people attacking the AWS team, at worst the “VP” who has to approve changes to the dashboard. That’s management and that “VP” is paid a lot.
Post reply on HN