Live data from Hacker News

Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

aws.amazon.com

121–130 of 410 posts

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#121

Earlier quoted context omitted.

I worked at Amazon. While my boss was on vacation I took over for him in the "Launch readiness" meeting for our team's component of our project. Basically, you go to this meeting with the big decision makers and business people once a week and tell them what your status is on deliverables. You are supposed to sum up your status as "Green/Yellow/Red" and then write (or update last week's document) to explain your stat…

> While my boss was on vacation I took over for him in the "Launch readiness" meeting....once a week Jeez, how many meetings did you go to, and how long was this person's vacation? I'm jelly of being allowed to take that much time off continuously.

You might be working at the wrong org? My colleagues routinely take weeks off at a time, sometimes more than a month to travel Europe, go scuba diving in French Polynesia, etc. Work to live, don’t live to work.

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#122

My favorite sentence: "Our networking clients have well tested request back-off behaviors that are designed to allow our systems to recover from these sorts of congestion events, but, a latent issue prevented these clients from adequately backing off during this event."

I saw pleeeeeenty of untested code at Amazon/AWS. Looking back it was almost like the most important services/code had the least amount of testing. While internal boondoggle projects (I worked on a couple) had complicated test plans and debates about coverage metrics.

This is almost always the case.

The most important services get the most attention from leaders who apply the most pressure, especially in the first ~2y of a fast-growing or high-potential product. So people skip tests.

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#123
post #57

Earlier quoted context omitted.

AWS has been getting a pass on their stability issues in us-east-1 for years now because it’s their “oldest” zone. Maybe they should invest in fixing it instead of inventing new services to sell.

I certainly wouldn't describe it as “a pass” given how commonly people joke about things like “friends don't let friends use us-east-1”. There's also a reporting bias: because many places only use us-east-1, you're more likely to hear about it even if it only affects a fraction of customers, and many of those companies blame AWS publicly because that's easier than admitting that they were only using one AZ, etc. Thes…

From the original outage thread:

“If you're having SLA problems I feel bad for you son I got two 9 problems cuz of us-east-1”

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#124

Does anyone know how often an AZ experiences an issue as compared to an entire region? AWS sells the redundancy of AZs pretty heavily, but it seems like a lot of the issues that happen end up being region-wide. I'm struggling to understand whether I should be replicating our service across regions or whether the AZ redundancy within a region is sufficient.

I've been naively setting up our distributed databases in separate AZs for a couple years now, paying, sometimes, thousands of dollars per month in data replication bandwidth egress fees. As far as I can remember I've never never seen an AZ go down, and the only region that has gone down has been us-east-1.

AZs definitely go down. It's usually due to a physical reason like fire or power issues.

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#125
post #85
post #74

Earlier quoted context omitted.

Some of us are in that camp and are looking at this outage and also pointing out that they continuously fail to accurately update their status dashboard in this and prior outages. Yes, doing what AWS does is hard, and yes outages /will/ happen, it is no knock on them that this outage occurred, what is a knock is that they haven't communicated honestly while the outage was ongoing.

They address that in the post, and between Twitter, HN and other places there wasn’t anyone legit questioning if something was actually broken. Contacts at AWS also all were very clear that yes something was going on and being investigated. This narrative that AWS was pretending nothing was wrong just wasn’t true based on what we saw.

I'm going to leave it at this: the dashboards at AWS aren't automated.

Say what you will, but I can automate a status dashboard in a couple days--yes, even at AWS scale.

No reason the dashboard should be green for hours while their engineers and support are aware things aren't working.

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#126

My company uses AWS. We had significant degradation for many of their APIs for over six hours, having a substantive impact on our business. The entire time their outage board was solid green. We were in touch with their support people and knew it was bad but were under NDA not to discuss it with anyone. Of course problems and outages are going to happen, but saying they have five nines (99.999) uptime as measured by…

> The entire time their outage board was solid green. We were in touch with their support people and knew it was bad but were under NDA not to discuss it with anyone. if ($pain > $gain) { move_your_shit_and_exit_aws(); } sub move_your_shit_and_exit_aws { printf("Dude. We have too much pain. Start moving\n"); printf("Yeah. That won't happen, so who cares\n"); exit(1); }

Moving your shit from AWS can be really expensive, depending on how much shit you have. If you're nice, GCP may subsidise - or even cover - the costs!

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#127

Does anyone know how often an AZ experiences an issue as compared to an entire region? AWS sells the redundancy of AZs pretty heavily, but it seems like a lot of the issues that happen end up being region-wide. I'm struggling to understand whether I should be replicating our service across regions or whether the AZ redundancy within a region is sufficient.

I've been naively setting up our distributed databases in separate AZs for a couple years now, paying, sometimes, thousands of dollars per month in data replication bandwidth egress fees. As far as I can remember I've never never seen an AZ go down, and the only region that has gone down has been us-east-1.

> I've never never seen an AZ go down, and the only region that has gone down has been us-east-1.

Doesn't the region going down mean that _all_ its AZs have gone down? Or is my mental model of this incorrect?

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#129
post #19

My company uses AWS. We had significant degradation for many of their APIs for over six hours, having a substantive impact on our business. The entire time their outage board was solid green. We were in touch with their support people and knew it was bad but were under NDA not to discuss it with anyone. Of course problems and outages are going to happen, but saying they have five nines (99.999) uptime as measured by…

I mean, not to defend them too strongly, but literally half of this post mortem is addressing the failure of the Service Dashboard. You can take it on bad faith, but they own up to the dashboard being completely useless during the incident.

So -- ctrl-f "Dash" only produces four results and it's hidden away in the bottom of the page. It's false to claim that even 20% of the post mortem is addressing the failure of the dashboard.

The problem is that the dashboard requires VP approval to be updated. Which is broken. The dashboard should be automatic. The dashboard should update before even a single member of the AWS team knows there's something wrong.

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#130
post #55

Queue the armchair infrastructure engineers. The reality is that there’s a handful of people in the world that can operate systems at this sheer scale and complexity and I have mad respect for those in that camp.

Isn't this the equivalent of "complaining about your meal in a restaurant, I'd like to see you do better."

The point of eating at a restaurant is that I can't/don't want to cook. Likewise, I use AWS because I want them to do the hard work and I'm willing to pay for it.

How does that abrogate my right to complain if it goes badly (regardless of whether I could/couldn't do it myself)?

Post reply on HN