I wish it contained actual detail and wasn’t couched in generalities.
Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
171–180 of 410 posts
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#172Earlier quoted context omitted.
The problem here sounds like lack of clarity over the meaning of the colours. In organisations with 100s of in-flight projects, it’s understandable that red is reserved for projects that are causing extremely serious issues right now. Otherwise, so many projects would be red, that you’d need a new colour.
How about orange? Didn't know there was a color shortage these days.
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#173Earlier quoted context omitted.
I mean, not to defend them too strongly, but literally half of this post mortem is addressing the failure of the Service Dashboard. You can take it on bad faith, but they own up to the dashboard being completely useless during the incident.
>You can take it on bad faith It's smart politics -- I don't blame them but I don't trust the dashboard either. There's established patterns now of the AWS dashboard being useless. If I want to check if Amazon is down I'm checking Twitter and HN. Not bad faith -- no faith.
Um, so you think straight-up lying is good politics?
Any 7-year old knows that telling a lie when you broke something makes you look better superficially, especially if you get away with it.
That does not mean that we should think it is a good idea to tell lies when you break things.
It sure as hell isn't smart politics in my book. It is straight-up disqualifying to do business with them. If they are not honest about the status or amount of service they are providing, how is that different than lying about your prices?
Would you go to a petrol station that posted $x.00/gallon, but only delivered 3 quarts for each gallon shown on the pump?
We're being shortchanged and lied to. Fascinating that you think it is good politics on their part.
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#174I was getting 503 "service unavailable" from STS during the outage most of the time I tried calling it.
I guess by "elevated latency", they mean from anyone with retry logic that would keep trying after many consecutive attempts?
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#175Noob question, but why does network infrastructure need dns? Why the full ipv6 address of the various components do not suffice to do business?
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#176Earlier quoted context omitted.
Off the top of my head, this is the third time they've had a major outage where they've been unable to properly update the status page. First we had the S3 outage, where the yellow and red icons were hosted in S3 and unable to be accessed. Second we had the Kinesis outage, which snowballed into a Cognito outage, so they were unable to login into the status page CMS. Now this. They "own up to it" in their postmortems,…
This challenge is not specific to Amazon. Being able to automatically detect system health is a non-trivial effort.
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#177Earlier quoted context omitted.
No, it was about an hour. We were aware from the very moment EC2 API error rates began to elevate, around 10:30 Eastern. By 11:30 the dashboard was updating. This timing is mentioned in the article, and it all happened in the middle of our workday on the east coast. The outage then continued for about 7 hours with SHD updates. I suspect we actually both agree on how long it took them to start updating, but I conclude…
At the large platform company where I work, our policy is if the customer reported the issue before our internal monitoring caught it, we have failed. Give 5 minutes for alerting lag, 10 minutes to evaluate the magnitude of impact, 10 minutes to craft the content and get it approved, 5 minutes to execute the update, adds up to 30 minutes end to end with healthy buffer at each step. 1 hour (52 minutes according to the…
They've discovered it right away, the Service Health Dashboard was not updated. source: link.
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#178> Customers accessing Amazon S3 and DynamoDB were not impacted by this event. We've seen plenty of S3 errors during that period. Kind of undermines credibility of this report.
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#179Earlier quoted context omitted.
Off the top of my head, this is the third time they've had a major outage where they've been unable to properly update the status page. First we had the S3 outage, where the yellow and red icons were hosted in S3 and unable to be accessed. Second we had the Kinesis outage, which snowballed into a Cognito outage, so they were unable to login into the status page CMS. Now this. They "own up to it" in their postmortems,…
This challenge is not specific to Amazon. Being able to automatically detect system health is a non-trivial effort.
Re: Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region
#180Problem is that I have to defend our own infrastructure real availability numbers vs cloud's fictional "five nines". It's a loosing game.
We have an environment we have access to for hosting webpages for one of the highest leaders in the whole Dept of Navy. This environment was DOWN (not "degrade availability" or "high latencies"), literally off of the Internet entirely, for CONSECUTIVE WEEKS earlier this year.
Completely incommunicado as well. It just happened to start working again one day. We collectively shrugged our shoulders and resumed updating our part of it.
This is an outlier example but even our normal sites I would classify as 1 "nine" of availability at best.