Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

891–900 of 1001 posts

Re: AWS us-east-1 outage

#891
post #874
post #245

Looks like they've acknowledged it on the status page now. https://status.aws.amazon.com/ > 8:22 AM PST We are investigating increased error rates for the AWS Management Console. > 8:26 AM PST We are experiencing API and console issues in the US-EAST-1 Region. We have identified root cause and we are actively working towards recovery. This issue is affecting the global console landing page, which is also hosted in US…

I like how 6 hours in: "Many services have already recovered".

Not even close for us.

Re: AWS us-east-1 outage

#893
post #878
post #877

Anyone else having trouble with their Alexa devices? Mine are acting really wonky, couldn’t listen to NPR in the morning really messed up my routine.

My Alexa-controlled lights didn’t turn on this evening.

Same for my smart plugs, this is infuriating to say the least. I hate having to re-enable skills in the Alexa app when these smart devices randomly decide to unlink.

Re: AWS us-east-1 outage

#894
Latest update:

"[4:35 PM PST] With the network device issues resolved, we are not working towards recovery of any impaired services. We will provide additional updates for impaired services within the appropriate entry in the Service Health Dashboard."

I guess they gave up.

Re: AWS us-east-1 outage

#895

Earlier quoted context omitted.

This feels like a discussion that could sorely use some numbers. What are good examples of >a small business running a few websites with a few million hits per month, it might be cheaper and easier to colocate a few servers and hire a few DevOps or old-school sysadmins to administer the infrastructure. and how often do they go down?

depends I guess, I am running on-prem workstation for our DWH. So far in 2 years it went down minutes at the time, when I decided to do so, because of hardware updates. I have no idea where this narrative came from, but usually hardware you have is very reliable and doesn't turn off every 15 minutes. Heck, I use old T430 for my home server and still it doesn't go down on completely random occasions (but thats very si…

But was it always accessible from the internet, and serving requests in an acceptable amount of time?

Re: AWS us-east-1 outage

#896

Latest update: "[4:35 PM PST] With the network device issues resolved, we are not working towards recovery of any impaired services. We will provide additional updates for impaired services within the appropriate entry in the Service Health Dashboard." I guess they gave up.

Hahaha, corrected to “now working” now.

Re: AWS us-east-1 outage

#897
post #863

Earlier quoted context omitted.

Utter lies on that page. Multiple services listed as green aren't working for me or my team.

Suggesting that when the status page sends a status request and hears no response—it defaults to green—hear no evil and see no evil —> report no evil Either way—overt lies or engineering incompetence—it’s disappointing!

Pretty low chance that the status page is automated, especially via health checks. I imagine it's a static asset updated by hand.

Re: AWS us-east-1 outage

#898

Earlier quoted context omitted.

I was not informed of his performance reviews. However, given the reception, his work in general, and the attitudes of the team, I cannot imagine this even came up. More likely the ability to improve routing to actually make YouTube cheaper in the end was I'm sure the ultimate positive result. This was also towards the end of the golden age of Google, when the percentage of top talent was a lot higher.

So on what basis is someone's performance reviewed, if such performance is omitted?

The entire point of blameless postmortems is acknowledging that the mere existence of an outage does not inherently reflect on the performance of the people involved. This allows you to instead focus on building resilient systems that avoid the possibility of accidental outages in the first place.

Re: AWS us-east-1 outage

#899

Latest update: "[4:35 PM PST] With the network device issues resolved, we are not working towards recovery of any impaired services. We will provide additional updates for impaired services within the appropriate entry in the Service Health Dashboard." I guess they gave up.

It was a typo. "Not" has changed to "now".

Re: AWS us-east-1 outage

#900
I'm pretty well versed in Kinesis / Firehose errors these days. I just got a new Firehose error I've never seen before:

> Firehose error: Slow down.

I wonder if they had to introduce a new exception / rate-limit to mediate this issue ...

Post reply on HN