Live data from Hacker News

Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

aws.amazon.com

391–400 of 410 posts

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#391

I’ve been running platform teams on aws now for 10 years, and working in aws for 13. For anyone looking for guidance on how to avoid this, here’s the advice I give startups I advise. First, if you can, avoid us-east-1. Yes, you’ll miss new features, but it’s also the least stable region. Second, go multi AZ for production workloads. Safety of your customer’s data is your ethical responsibility. Protect it, back it up…

> Third, you’re gonna go down when the cloud goes down. Not necessarily. You just need to not be stuck with a single cloud provider. The likelihood of more than one availability zone going down on a single cloud provider is not that low in practice. Especially when the problem is a software bug. The likelihood of AWS, Azure, and OVH going down at the same time is low. So if you need to stay online if AWS fail, don't…

Complex systems are expensive to operate, in many ways.

The more complexity you build into your own systems on top of the providers you depend on, the more likely you are to shoot yourself in the foot when you run into complexity issues that you’ve never seen before.

And the times that is most likely to happen is when one of your complex service providers goes down.

If the kind of thing you’re talking about could be feasibly done, then Netflix would have already done it. The fact that Netflix hasn’t solved this problem is a strong indicator that piling more proprietary complexity on top of all the vendor complexity you inherit from using a given service, well that’s a really hard problem in and of itself.

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#392
post #85

Earlier quoted context omitted.

They address that in the post, and between Twitter, HN and other places there wasn’t anyone legit questioning if something was actually broken. Contacts at AWS also all were very clear that yes something was going on and being investigated. This narrative that AWS was pretending nothing was wrong just wasn’t true based on what we saw.

I'm going to leave it at this: the dashboards at AWS aren't automated. Say what you will, but I can automate a status dashboard in a couple days--yes, even at AWS scale. No reason the dashboard should be green for hours while their engineers and support are aware things aren't working.

Uh, no. You can’t. If you could, then you would already have been hired and you would have already solved this problem.

What you can do at what you think is AWS scale has no bearing on what you could actually do at real AWS scale.

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#393

Earlier quoted context omitted.

> I wish we would just throw up a generic "Shit's Fucked Up. We Don't Know Why Yet, But We're Working On It" message. I gotta say, the implication that you can't register an outage until you know why it happened is pretty damning. The status page is where we look to see if services are effected, if that information can't be shared there until you understand the cause, that's very broken. The AWS status page has becom…

Can you please help me understand why you, and everyone else, are so passionate about the status page? I get that it not being updated is an annoyance, but I cannot figure out why it is the single most discussed thing about this whole event. I mean, entire services were out for almost an entire day, and if you read HN threads it would seem that nobody even cares about lost revenue/productivity, downtime, etc. The vas…

I care about status pages, because when something breaks upstream I need to know whether it's an issue I need to report, and if there's additional problems related to the outage I need to look out for, or workarounds I can deploy. If I find out anything that might help me narrow down the ETA for a fix, that's bonus fries.

I don't gripe about it on HN, but it is generally a disappointment to me when I stumble upon something that looks like a significant outage but a company is making no indication that they've seen it and are working on it (or waiting for something upstream of them, as sometimes happens).

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#394

Earlier quoted context omitted.

> a VP must sign off on changing status pages, which is... backwards to say the least. I think most people's experience with "VP's" makes them not realize what AWS VP's do. VP's here are not sitting in an executive lounge wining and dining customers, chomping on cigars and telling minions to "Call me when the data center is back up and running again!" They are on the tech call, working with the engineers, evaluating…

It doesn’t matter what the VPs are doing, that misses the point. Every minute you know there is a problem and you haven’t at least put up a “degraded” status, you’re lying to your customers. It was on the top of HN for an hour before anything changed, and then it was still downplayed, which is insane.

The vps are busy getting status and making busy work for working level folks. The guy defending the vp is probably a vp or a vp helper.

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#395

Earlier quoted context omitted.

Saying "S3 is down" can mean anything. Our S3 buckets that served static web content stayed up no problem. The API was down though. But for the purposes of whether my organization cares I'm gonna say it was "up".

Most of what they actually said via the manual human-language status updates was "Service X is seeing elevated error rates". While there are still decisions to be made in how you monitor errors and what sorts of elevated rates merit an alert -- I would bet that AWS has internally-facing systems that can display service health in this way based on automated monitoring of error rates (as well as other things). Because…

As mentioned in the article, internal metrics were fubar most of the day.

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#396

Earlier quoted context omitted.

> Third, you’re gonna go down when the cloud goes down. Not much use getting overly bent out of shape. Ugh. I have a hard time with this one. Back in the day, EBS had some really awful failures and degradations. Building a greenfield stack that specifically avoided EBS and stayed up when everyone else was down during another mass EBS failure felt marvelous. It was an obvious avoidable hazard. It doesn't mean "avoid E…

I hear you. I didn’t use EBS for five years after the great outage in, what was it, 2011? At this point, it’s reliable enough that even if it were to go down, it’s more safe than not using it. I’d put EBS in the pantheon of “core” services I never mind using these days.

Yup, 2011. That's the one. One of those US presidential campaigns stayed up throughout because of EBS-phobia.

Geez. We have decades-old cloud war stories now? I suddenly feel really old.

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#397
post #301

Earlier quoted context omitted.

Servers are not hard if you have a dedicated person (long time ago known as Systemadminstrator), and fun fact...it's sometimes even much cheaper and more reliable then having everything in the "cloud". Personally i am a believer in mixed environments, public webservers etc in the "cloud", locally used systems and backup "in house" with a second location (both in Data-centers or at least one), and no, i don't talk abo…

You could have a staff of a million-plus people and stuff could still go sideways. Hint — it did.

Wow Sherlock....no BS? Look at the title of this article...

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#398
post #353

Earlier quoted context omitted.

99% of businesses don't need 24/7 but two are the bare minimum (a admin, and a dev or admin)

That’s like saying that you’ve got a hard drive so you could be your own DropBox. You vastly underestimate the amount of resources required. And what those resources cost.

That's why i wrote "not the next google, but the other 99%", and now that HN-guy comes around and compares it with dropbox....

99% of company's are hairdressers, lawyers, builders, insurance etc.....and NOT the next google/dropbox/facebook...are you so far away from reality?

I know how much it costs, because that's exactly what i do since ~19 years.

Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region

#400

Earlier quoted context omitted.

Interesting. Just wondering if your guys have a dedicated DBA?

Not sure why I got down voted for an honest question. Most start-ups are founders, developers, sales and marketing. Dedicated infrastructure, network and database specialists don't get factored in because "smart CS graduates can figure that stuff out". I've worked at companies who held onto that false notion way too long and almost lost everything as a result ("company extinction event", like losing a lot of customer…

I am always amazed at how little my software dev spouse understands about infrastructure, basic networking troubleshooting is beyond her. She is a great dev, but a terrible at ops. Fortunately she is at a large company with lots of devs, sysadmins and SREs.
Post reply on HN