Live data from Hacker News

Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

aws.amazon.com

331–340 of 410 posts

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#331

Earlier quoted context omitted.

> I wish we would just throw up a generic "Shit's Fucked Up. We Don't Know Why Yet, But We're Working On It" message. I think that's the crux of the matter? AWS seems to now have a reputation for ignoring issues that are easily observable by customers, and by the time any update shows up, it's way too late. Whether VPs make this decision or not is irrelevant. If this becomes a known pattern (and I think it has), then…

Saying "S3 is down" can mean anything. Our S3 buckets that served static web content stayed up no problem. The API was down though. But for the purposes of whether my organization cares I'm gonna say it was "up".

Who cares if it worked for your usecase?

Being unable to store objects in an object store means that it’s broken.

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#332

Earlier quoted context omitted.

I’m not all that angry over the situation but more disappointed that we’ve all collectively handed the keys over to AWS because “servers are hard”. Yeh they are but it’s not like locking ourselves into one vendor with flaky docs and a black box of bugs is any better, at least when your own servers go down it’s on you and you don’t take out half of North America.

You can either pay a dedicated team to manage your on prem solution, go multi cloud, or simply go multi region on aws. My company was not affected by this outage because we are multi region. Cheapest and quickest option if you want to have at least some fault tolerance.

So was mine, but we couldn't log in

But yes, having services resilient to a single point of failure is essential. AWS is a SPOF.

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#333
post #278

Complex systems are really really hard. I'm not a big fan of seeing all these folks bash AWS for this, and not really understanding the complexity or nastiness of situations like this. Running the kind of services they do for the kind of customers, this is a VERY hard problem. We ran into a very similar issue, but at the database layer in our company literally 2 weeks ago, where connections to our MySQL exploded and…

> I'm not a big fan of seeing all these folks bash AWS for this,

The disdain I saw was towards those claiming that all you need is AWS, that AWS never goes down, and don't bother planning for what happens when AWS goes down.

AWS is an amazing accomplishment, but it's still a single point of failure. If you are a company relying on a single supplier and you don't have any backup plans for that supplier being unavailable, that is ridiculous and worthy of laughter.

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#334

Earlier quoted context omitted.

I think most of the outrage is not because "it happened" but because AWS is saying things like "S3 was unaffected" when the anecdotal experience of many in this thread suggests the opposite. That and the apparent policy that a VP must sign off on changing status pages, which is... backwards to say the least.

> a VP must sign off on changing status pages, which is... backwards to say the least. I think most people's experience with "VP's" makes them not realize what AWS VP's do. VP's here are not sitting in an executive lounge wining and dining customers, chomping on cigars and telling minions to "Call me when the data center is back up and running again!" They are on the tech call, working with the engineers, evaluating…

There should be no one to sign off anything. Your status page should be updated automatically, not manually!

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#335

I’ve been running platform teams on aws now for 10 years, and working in aws for 13. For anyone looking for guidance on how to avoid this, here’s the advice I give startups I advise. First, if you can, avoid us-east-1. Yes, you’ll miss new features, but it’s also the least stable region. Second, go multi AZ for production workloads. Safety of your customer’s data is your ethical responsibility. Protect it, back it up…

> Safety of your customer’s data is your ethical responsibility. Protect it, back it up, keep it as generally available as is reasonable.

> Third, you’re gonna go down when the cloud goes down. Not much use getting overly bent out of shape.

“Whoops, our provider is down, sorry!” is not taking responsibility with customer data at all.

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#336

I am grateful to AWS for this report. Not sure if any AWS support staff are monitoring this thread, but the article said: > Customers also experienced login failures to the AWS Console in the impacted region during the event. All our AWS instances / resources are in EU/UK availability zones, and yet we couldn't access our console either. Thankfully none of our instances were affected by the outage, but our inability…

They posted on the status page to try using the alternate region endpoints like us-west.console.Amazon.com (I think) at the time, but not sure if it was a true fix.

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#337
post #49

> Customers accessing Amazon S3 and DynamoDB were not impacted by this event. We've seen plenty of S3 errors during that period. Kind of undermines credibility of this report.

Yeah > For example, while running EC2 instances were unaffected by this event That ignores the 3-5% drop in traffic I saw in us-east-1 on EC2 instances that only talk to peers on the Internet with TCP/IP during this event.

How are you measuring this? Remember that cloudwatch was apparently also losing metrics, so aggregating CW metrics might show that kind of drop.

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#338
post #278

Complex systems are really really hard. I'm not a big fan of seeing all these folks bash AWS for this, and not really understanding the complexity or nastiness of situations like this. Running the kind of services they do for the kind of customers, this is a VERY hard problem. We ran into a very similar issue, but at the database layer in our company literally 2 weeks ago, where connections to our MySQL exploded and…

> Outrage is the easy response. Empathy and learning is the valuable one.

I'm outraged that AWS, as a company policy, continues to lie about the status of their systems during outages, making it hard for me to communicate to my stakeholders.

Empathy? For AWS? AWS is part a mega corporation that is closing in on 2 TRILLION dollars in market cap. It's not a person. I can empathize with individuals who work for AWS but it's weird to ask us to have empathy for a massive faceless, ruthless, relentless, multinational juggernaut.

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#339
There are a lot of comments in here that boil down to "could you do infrastructure better?"

No, absolutely not. That's why I'm on AWS.

But what we are all ACTUALLY complaining about is ongoing lack of transparent and honest communications during outages and, clearly, in their postmortems.

Honest communications? Yeah, I'm pretty sure I could do that much better than AWS.

Post reply on HN