Earlier quoted context omitted.
> I cant think of one. I can. "S3 is unavailable because X, Y, and Z services are unavailable." A graph of dependencies between services is surely known to AWS; if not, they ought to create one post-haste. Trying to externalize Amazon's internal AWS politicking over which service is down is unproductive to the customers who check the dashboard and see that their service ought to be up, but... well, it isn't? Because…
Yes, I can envision a (simplified) AWS X-Ray dashboard showing the relationships between the systems and the performance of each one. Then we could see at a glance what was going on. Almost anything is better than that wall of text, tiny status images, and RSS feeds.
Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
351–360 of 410 posts
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#352Earlier quoted context omitted.
I mean, not to defend them too strongly, but literally half of this post mortem is addressing the failure of the Service Dashboard. You can take it on bad faith, but they own up to the dashboard being completely useless during the incident.
Off the top of my head, this is the third time they've had a major outage where they've been unable to properly update the status page. First we had the S3 outage, where the yellow and red icons were hosted in S3 and unable to be accessed. Second we had the Kinesis outage, which snowballed into a Cognito outage, so they were unable to login into the status page CMS. Now this. They "own up to it" in their postmortems,…
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#353Earlier quoted context omitted.
Servers are not hard if you have a dedicated person (long time ago known as Systemadminstrator), and fun fact...it's sometimes even much cheaper and more reliable then having everything in the "cloud". Personally i am a believer in mixed environments, public webservers etc in the "cloud", locally used systems and backup "in house" with a second location (both in Data-centers or at least one), and no, i don't talk abo…
Not one person, at least four people to run stuff 24/7.
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#354Earlier quoted context omitted.
Yeah > For example, while running EC2 instances were unaffected by this event That ignores the 3-5% drop in traffic I saw in us-east-1 on EC2 instances that only talk to peers on the Internet with TCP/IP during this event.
How are you measuring this? Remember that cloudwatch was apparently also losing metrics, so aggregating CW metrics might show that kind of drop.
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#355Earlier quoted context omitted.
If I recall there was a point in time where the control panel for all regions was in us-east-1. I seem to recall an outrage where the other regions were up, but you couldn’t change any resources because the management api was down in us-east-1
This was our exact experience with this outage. Literally all our AWS resources are in EU/UK regions - and they all continued functioning just fine - but we couldn't sign in to our AWS console to manage said resources. Thankfully the outage didn't impact our production systems at all, but our inability to access said console was quite alarming to say the least.
It would probably be clearer that they exist if the console redirected to the regional URL when you switched regions.
STS, S3, etc have regional endpoints too that have continued to work when us-east-1 has been broken in the past and the various AWS clients can be configured to use them, which they also sadly don't tend to do by default.
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#356Earlier quoted context omitted.
>>It's smart politics -- I don't blame them Um, so you think straight-up lying is good politics? Any 7-year old knows that telling a lie when you broke something makes you look better superficially, especially if you get away with it. That does not mean that we should think it is a good idea to tell lies when you break things. It sure as hell isn't smart politics in my book. It is straight-up disqualifying to do busi…
You don’t know what you’re talking about. AWS spends a lot of time thinking about this problem in service to their customers. How do you reduce the status of millions of machines, the software they run, and the interconnected-ness of those systems to a single graphical indicator? It would be dumb and useless to turn something red every single time anything had a problem. Literally there are hundreds of things broken…
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#357Complex systems are really really hard. I'm not a big fan of seeing all these folks bash AWS for this, and not really understanding the complexity or nastiness of situations like this. Running the kind of services they do for the kind of customers, this is a VERY hard problem. We ran into a very similar issue, but at the database layer in our company literally 2 weeks ago, where connections to our MySQL exploded and…
I’m not all that angry over the situation but more disappointed that we’ve all collectively handed the keys over to AWS because “servers are hard”. Yeh they are but it’s not like locking ourselves into one vendor with flaky docs and a black box of bugs is any better, at least when your own servers go down it’s on you and you don’t take out half of North America.
A couple years ago all our services at our data center just vanished. I call the data center and they start creating a ticket. "Can you tell me if there is a data center outage?" "We are currently investigating and I don't have any information I can give you." "Listen, if this is a problem isolated to our cabinet, I need to get in the car. I'm trying to decide if I need to drive 60 miles in a blizzard."
That facility has been pretty good to us over a decade, but they were frustratingly tight-lipped about an entire room of the facility losing power because one of their power feeder lines was down.
Could AWS improve? Yes. Does avoiding AWS solve these sorts of problems? No.
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#358Earlier quoted context omitted.
> a VP must sign off on changing status pages, which is... backwards to say the least. I think most people's experience with "VP's" makes them not realize what AWS VP's do. VP's here are not sitting in an executive lounge wining and dining customers, chomping on cigars and telling minions to "Call me when the data center is back up and running again!" They are on the tech call, working with the engineers, evaluating…
> I wish we would just throw up a generic "Shit's Fucked Up. We Don't Know Why Yet, But We're Working On It" message. I gotta say, the implication that you can't register an outage until you know why it happened is pretty damning. The status page is where we look to see if services are effected, if that information can't be shared there until you understand the cause, that's very broken. The AWS status page has becom…
I get that it not being updated is an annoyance, but I cannot figure out why it is the single most discussed thing about this whole event. I mean, entire services were out for almost an entire day, and if you read HN threads it would seem that nobody even cares about lost revenue/productivity, downtime, etc. The vast majority of comments in all of the outage threads are screaming about how the SHD lied.
In my entire career of consulting across many companies and many different technology platforms, never once have I seen or heard of anyone even looking at a status page outside of HN. I'm not exaggerating. Even over the last 5 years when I've been doing cloud consulting, nobody I've worked with has cared at all about the cloud provider's status pages. The only time I see it brought up is on HN, and when it gets brought up on HN it's discussed with more fervor than most other topics, even the outage itself.
In my real life (non-HN) experience, when an outage happens, teams ask each other "hey, you seeing problems with this service?" "yea, I am too, heard maybe it's an outage" "weird, guess I'll try again later" and go get a coffee. In particularly bad situations, they might check the news or ask me if I'm aware of any outage. Either way, we just... go on with our lives? I've never needed, nor have I ever seen people need, a status page to inform them that things aren't working correctly, but if you read HN you would get the impression that entire companies of developers are completely paralyzed unless the status page flips from green to red. Why? I would even go as far to say that if you need a third party's SHD to tell you if things aren't working right, then you're probably doing something wrong.
Seriously, what gives? Is all this just because people love hating on Amazon and the SHD is an easy target? Because that's what it seems like.
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#359Earlier quoted context omitted.
> I wish we would just throw up a generic "Shit's Fucked Up. We Don't Know Why Yet, But We're Working On It" message. I gotta say, the implication that you can't register an outage until you know why it happened is pretty damning. The status page is where we look to see if services are effected, if that information can't be shared there until you understand the cause, that's very broken. The AWS status page has becom…
Can you please help me understand why you, and everyone else, are so passionate about the status page? I get that it not being updated is an annoyance, but I cannot figure out why it is the single most discussed thing about this whole event. I mean, entire services were out for almost an entire day, and if you read HN threads it would seem that nobody even cares about lost revenue/productivity, downtime, etc. The vas…
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#360Earlier quoted context omitted.
>>It's smart politics -- I don't blame them Um, so you think straight-up lying is good politics? Any 7-year old knows that telling a lie when you broke something makes you look better superficially, especially if you get away with it. That does not mean that we should think it is a good idea to tell lies when you break things. It sure as hell isn't smart politics in my book. It is straight-up disqualifying to do busi…
You don’t know what you’re talking about. AWS spends a lot of time thinking about this problem in service to their customers. How do you reduce the status of millions of machines, the software they run, and the interconnected-ness of those systems to a single graphical indicator? It would be dumb and useless to turn something red every single time anything had a problem. Literally there are hundreds of things broken…
There's a limitless variety of options, and multiple books written about it. I can recommend the series "The Visual Display of Quantitative Information" by Edward Tufte, for starters.
>> Literally there are hundreds of things broken every minute of every day. On-call engineers are working around the clock...
Of course there are, so a single R/Y/G indicator is obviously a bad choice.
Again, they could at any time easily choose a better way to display this information, graphs, heatmaps, whatever.
More importantly, the one thing that should NOT be chosen is A) to have a human in the loop of displaying status, as this inserts both delay and errors.
Worse yet, to make it so that it is a VP-level decision, as if it were a $1million+ purchase, and then to set the policy to keep it green when half a continent is down... ummm that is WAAAYYY past any question of "threshold" - it is a premeditated, designed-in, systemic lie.
>>You don’t know what you’re talking about. Look in the mirror, dude. While I haven't worked inside AWS, I have worked in complex network software systems and well understand the issues of thousands of HW/SW components in multiple states. More importantly, perhaps it's my philosophy degree, but I can sort out WHEN (e.g., here) the problem is at another level altogether. It is not the complexity of the system that is the problem, it is the MANAGEMENT decision to systematically lie about that complexity. Worse yet, it looks like those lies on an everyday basis are what goes into their claims of "99.99+% uptime!!" evidently false. The problem is at the forest level, and you don't even want to look at the trees because you're stuck in the underbrush telling everyone else they are clueless.