Live data from Hacker News

AWS's us-east-1 region is experiencing issues

health.aws.amazon.com

91–100 of 168 posts

Re: AWS's us-east-1 region is experiencing issues

#91

Earlier quoted context omitted.

Or perhaps triaging, root-causing, and fixing the issue is the highest-order bit?

Different people have different responsibilities. At Amazon scale, the comms and people doing a deep dive to fix stuff will not be the same.

I'd be totally fine just having alerts and metrics driving the status page. Why involve a human at all? They just get emotional.

(I have a data-driven status page for my personal website. If Oh Dear decides my website is down, the status page gets automatically updated. Obviously nobody is ever going to visit status.jrock.us if they are trying to read an article on my blog and it doesn't load, but hey at least I can say I did it.)

Re: AWS's us-east-1 region is experiencing issues

#92
post #80

Earlier quoted context omitted.

A handful of large-traffic sites have recently, and relatively suddenly, started blocking traffic from a large region. That's a major change in flow.

Could you be more specific?

Russia / sanctions, I'm guessing.

Re: AWS's us-east-1 region is experiencing issues

#93

Earlier quoted context omitted.

> What's causing people to believe that the latest round of attrition is any different? The Great Recession The Great Resignation The Great Dying (due to COVID-19)

> The Great Dying (due to COVID-19) Repeating this wise comment: https://news.ycombinator.com/item?id=23769427 The COVID death counts are hopelessly over-counted. This is why there's a cottage industry of people pointing out things like "COVID deaths" which mysteriously also suffered from being murdered, or drug overdoses, or undiagnosed leukaemia. Then you get into the problem of care homes being authorised to repor…

[deleted]

Re: AWS's us-east-1 region is experiencing issues

#94
post #72

Earlier quoted context omitted.

I’ll be honest. My external impression of Amazon and Google could not be more distant in this regard. Google doesn’t have nearly as hard a time retaining good people as Amazon does.

From within AWS, it just feels like we push people too hard. Service teams are too small relative to their goals, sales teams have unrealistic growth targets and also double as support for needy/incompetent customers, and professional services have billable hour requirements which are as high as any major consultancy and with additional pre-sales support expectations. Strict adherence to the "hiring bar" means we fai…

>"Strict adherence to the "hiring bar" means we fail to bring in good people who aren't desperate enough to act out the cultish LP dance during their interview."

What is the "cultish LP dance" here that is weeding good people out?

>"My team is hiring for 2-3 people and we are being buried alive without that growth happening sooner - but I can't in good conscience recommend this place to anyone I respect or like."

I appreciate your candor. Are you in a dev role or are you on the SRE side? Is your description true across pretty much all teams/services then?

Re: AWS's us-east-1 region is experiencing issues

#95

Earlier quoted context omitted.

> What's causing people to believe that the latest round of attrition is any different? The Great Recession The Great Resignation The Great Dying (due to COVID-19)

> The Great Dying (due to COVID-19) Repeating this wise comment: https://news.ycombinator.com/item?id=23769427 The COVID death counts are hopelessly over-counted. This is why there's a cottage industry of people pointing out things like "COVID deaths" which mysteriously also suffered from being murdered, or drug overdoses, or undiagnosed leukaemia. Then you get into the problem of care homes being authorised to repor…

How about undercounting?

There’s very likely to be severe undercounting of COVID-19 due to the same reasons that crimes of victimization are underreported: shame.

Not to mention the swaths of homeless and disabled people that probably didn’t get counted.

Re: AWS's us-east-1 region is experiencing issues

#96

Earlier quoted context omitted.

I would strongly urge not using us-east-1 -- of all the regions we're in, it's by far the most problematic. Use us-east-2 if you need good latency to the East Coast.

Not sure if it's still the case, but when I was there us-east-1 was a SPOF for some services world wide . I think if dynamodb went down in the region it was a big, big issue.

The only SPOF of failure I know of for us-east-1 today is the control plane for Route53 - it's distributed and DNS queries will continue to work when us-east-1 is down (including health check based failover), but you can't make any DNS changes when us-east-1 is down.

Re: AWS's us-east-1 region is experiencing issues

#97

Earlier quoted context omitted.

From within AWS, it just feels like we push people too hard. Service teams are too small relative to their goals, sales teams have unrealistic growth targets and also double as support for needy/incompetent customers, and professional services have billable hour requirements which are as high as any major consultancy and with additional pre-sales support expectations. Strict adherence to the "hiring bar" means we fai…

>"Strict adherence to the "hiring bar" means we fail to bring in good people who aren't desperate enough to act out the cultish LP dance during their interview." What is the "cultish LP dance" here that is weeding good people out? >"My team is hiring for 2-3 people and we are being buried alive without that growth happening sooner - but I can't in good conscience recommend this place to anyone I respect or like." I a…

> What is the "cultish LP dance" here that is weeding good people out?

The "culture fit" interview process focuses on leadership principles, so lots of questions like " tell me about a time when you went above and beyond for a customer". Being yourself will get you nowhere, you need to research the questions and the script that is expected of you.

> What service does your team work on?

I'm a partner-focused SA, so not a developer and not aligned to a particular service.

Re: AWS's us-east-1 region is experiencing issues

#98
post #77

Maybe the reason AWS keeps going down is because they run all their stuff on-prem...

I'm not sure if gcloud or azure would help. I run two servers on hetzner which is way cheaper than azure/gcloud they would be better off there.

They might benefit from migrating to the Azure cloud. I’ve heard that some of the Windows servers actually run faster than some of the Linux servers on Azure.

Re: AWS's us-east-1 region is experiencing issues

#99

Earlier quoted context omitted.

Different people have different responsibilities. At Amazon scale, the comms and people doing a deep dive to fix stuff will not be the same.

I'd be totally fine just having alerts and metrics driving the status page. Why involve a human at all? They just get emotional. (I have a data-driven status page for my personal website. If Oh Dear decides my website is down, the status page gets automatically updated. Obviously nobody is ever going to visit status.jrock.us if they are trying to read an article on my blog and it doesn't load, but hey at least I can…

> Why involve a human at all?

To make a judgement call on whether the issue is severe enough to warrant the legal/financial risk of admitting your service is broken, potentially breaking customer SLAs.

Re: AWS's us-east-1 region is experiencing issues

#100
post #8

This is why you are strongly urged not to rely on one region or AZ.

Not always possible - Australia (currently) only has one availability zone and if you're in a regulated industry (banking or government stuff) they require data to be in Australia.
Post reply on HN