Live data from Hacker News

AWS's us-east-1 region is experiencing issues

health.aws.amazon.com

141–150 of 168 posts

Re: AWS's us-east-1 region is experiencing issues

#141

Earlier quoted context omitted.

> What is the "cultish LP dance" here that is weeding good people out? The "culture fit" interview process focuses on leadership principles, so lots of questions like " tell me about a time when you went above and beyond for a customer". Being yourself will get you nowhere, you need to research the questions and the script that is expected of you. > What service does your team work on? I'm a partner-focused SA, so no…

There is nothing cultish about the way Amazon interviews. If you are a good engineer with a relevant background who can speak english you will have no problem passing these interviews. I'm not sure why you are framing it as cultish.

I interviewed with AWS about 7 months ago and got an offer. I had multiple recruiters emphasize the LP stuff. I prepared, and there’s no way I would have passed without that preparation.

In my experience, they also give the toughest programming questions. It is a lot of prep overall.

Re: AWS's us-east-1 region is experiencing issues

#142

Earlier quoted context omitted.

There is nothing cultish about the way Amazon interviews. If you are a good engineer with a relevant background who can speak english you will have no problem passing these interviews. I'm not sure why you are framing it as cultish.

Native English speaker here. I was denied for unspecified reasons a couple years ago. At the time it was perplexing, as I had pretty good answers for all their questions. Now I work at one of the slightly more sane FANMAG (to include $ms) companies. Pretty sure I dodged a bullet, maybe the engineering manager spared me because he liked me more than I realized.

I would be curious to know which one is considered more sane these days. I feel like I've heard enough negative things about the culture at all of them at this point.

Re: AWS's us-east-1 region is experiencing issues

#143
post #54
post #26

Earlier quoted context omitted.

Just like the tech priests in Warhammer 40k, keeping occult old engineering, thatno one could build anymore, running

So today I find out my job title is tech priest. I was happy with necromancer before. Does it come with a pay rise?

Not at all, but a status increase for sure.

Re: AWS's us-east-1 region is experiencing issues

#144

Earlier quoted context omitted.

>"Strict adherence to the "hiring bar" means we fail to bring in good people who aren't desperate enough to act out the cultish LP dance during their interview." What is the "cultish LP dance" here that is weeding good people out? >"My team is hiring for 2-3 people and we are being buried alive without that growth happening sooner - but I can't in good conscience recommend this place to anyone I respect or like." I a…

> What is the "cultish LP dance" here that is weeding good people out? The "culture fit" interview process focuses on leadership principles, so lots of questions like " tell me about a time when you went above and beyond for a customer". Being yourself will get you nowhere, you need to research the questions and the script that is expected of you. > What service does your team work on? I'm a partner-focused SA, so no…

I interviewed with AWS about a year ago. Knocked the programming/system design interviews out of the park, but it was clear the LP interviewer simply didn’t believe I was being truthful about the example I gave (included a period where my team had no direct manager, which is abnormal). He also had a programming question for me but we didn’t have time for him to explain it.

No offer, recruiter emphasized that they were halving the “cool off” period for me (so I could interview again soon), and maybe they do this for everyone, but it’s clear there was one interview making the difference here. Interesting that this is apparently a common problem.

Re: AWS's us-east-1 region is experiencing issues

#145

Earlier quoted context omitted.

Different people have different responsibilities. At Amazon scale, the comms and people doing a deep dive to fix stuff will not be the same.

I'd be totally fine just having alerts and metrics driving the status page. Why involve a human at all? They just get emotional. (I have a data-driven status page for my personal website. If Oh Dear decides my website is down, the status page gets automatically updated. Obviously nobody is ever going to visit status.jrock.us if they are trying to read an article on my blog and it doesn't load, but hey at least I can…

Corey Quinn has a great blog post[0] on why status pages are hard, especially for large organizations with lots of products.

[0] https://www.lastweekinaws.com/blog/status-paging-you/

Re: AWS's us-east-1 region is experiencing issues

#146
post #99

Earlier quoted context omitted.

I'd be totally fine just having alerts and metrics driving the status page. Why involve a human at all? They just get emotional. (I have a data-driven status page for my personal website. If Oh Dear decides my website is down, the status page gets automatically updated. Obviously nobody is ever going to visit status.jrock.us if they are trying to read an article on my blog and it doesn't load, but hey at least I can…

> Why involve a human at all? To make a judgement call on whether the issue is severe enough to warrant the legal/financial risk of admitting your service is broken, potentially breaking customer SLAs.

If you're just going to lie, why have an SLA at all? It's like doing a clinical trial for a drug; a bunch of your patients die and you say "well they were going to die anyway it has nothing to do with the drug." If it's one person, maybe you can get away with that. When it's everyone in the experimental group, people start to wonder.

I have two arguments in favor of honest SLAs. One is, if customers expect that something is down, it can give them a piece of data with which to make their mitigation decisions. "A lot of our services are returning errors", check the auto-updated status page, "there may be an issue with network routes between AZs A and C". Now you know to drain out of those zones. If the status page says "there are no problems", now you spend hours debugging the issue, costing yourself far more money in your time than you spend on your cloud infrastructure in the first place. If having an SLA is the cause of that, it would be financially in your favor to not have the SLA at all. The SLA bounds your losses to what you pay for cloud resources, but your losses can actually be much higher; lost revenue, lost time troubleshooting someone else's problem, etc.

The second is, SLA violations are what justify reliability engineering efforts. If you lose $1,000,000 a year to SLA violations, and you hire an SRE for $500,000 a year to reduce SLA violations by 75%, then you just made a profit of $250,000. If you say "nah there were no outages", then you're flushing that $500,000 a year down the toilet and should fire anyone working on reliability. That is obviously not healthy; the financial aspect keeps you honest and accountable.

All of this gets very difficult when you are planning your own SLAs. If everyone is lying to you, you have no choice but to lie to your customers. You can multiply together all the 99.5% SLAs of the services you depend on and give your customers a guarantee of 95%, but if the 99.5% you're quoted is actually 89.4%, then you can't actually meet your guarantees. AWS can afford to lie to their customers (and Congress apparently) without consequences. But you, small startup, can't. Your customers are going to notice, and they were already taking a chance going with you instead of some big company. This is hard cycle to get out of. People don't want to lie, but they become liars because the rest of the industry is lying.

Finally, I just want to say I don't even care about the financial aspect, really. The 5 figures spent on cloud expenses are nothing compared to the nights and weekends your team loses to debugging someone else's problem. They could have spent the time with their families or hobbies if the cloud provider just said "yup, it's all broken, we'll fix it by tomorrow". Instead, they stay up late into the night looking for the root cause, finding it 4 hours later, and still not being able to do anything except open a support ticket answered by someone who has to lie to preserve the SLA. They'll never get those hours back! And they turned out to be a complete waste of time.

It's a disaster, and I don't care if it's a hard problem. We, as an industry, shouldn't stand for it.

Re: AWS's us-east-1 region is experiencing issues

#147

I can't help but wonder, with the increases in attrition across the industry, are we hitting some kind of tipping point where the institutional knowledge in these massive tech corporations is disappearing? Mistakes happen all the time but when all the people who intimately know how these systems work leave for other opportunities, disasters are bound to happen more and more.

Yes, absolutely. Within my own org of ~50 people, 15% have resigned/contracts ending during Q1 (after 15% in Q4). Of the remaining 85%.. 20% have been around since before COVID / 65% joined during COVID. Of our senior engineers & team leads, 70% have joined in last 6-9 months. Only 3 full time senior engineers with 2 years or more tenure. We've grown during COVID but we've also just burned through people. Turnover ha…

Newest member of my team has been in the company for 6 years and on my team for 4. I was in the local pub the other day and there was a retirement do for someone who had been here for 35 years, which certainly isn't exceptional (40 years is more noteworthy, and I've known a couple of people who made it to 50 years)

Re: AWS's us-east-1 region is experiencing issues

#148

I can't help but wonder, with the increases in attrition across the industry, are we hitting some kind of tipping point where the institutional knowledge in these massive tech corporations is disappearing? Mistakes happen all the time but when all the people who intimately know how these systems work leave for other opportunities, disasters are bound to happen more and more.

Resolved in 7 mins. Can you do better?

My monitoring doesn't remember the last time we had a service outage lasting 7 seconds, let alone 7 minutes.

Re: AWS's us-east-1 region is experiencing issues

#149

Earlier quoted context omitted.

> What is the "cultish LP dance" here that is weeding good people out? The "culture fit" interview process focuses on leadership principles, so lots of questions like " tell me about a time when you went above and beyond for a customer". Being yourself will get you nowhere, you need to research the questions and the script that is expected of you. > What service does your team work on? I'm a partner-focused SA, so no…

As some dude from Baltimore who has just picked up a second gig there. It seemed to me that these are normal questions asked of you in most interviews. In my daytime position at the first gig I have interviewed technicians and have asked similar questions of them. It's not about culture fit. It's about finding their answer to basic customer service questions. It would make sense that a customer obsessed company would…

They are normal-ish questions and I can see where the STAR guidance is better than getting no help at all - I just think the process is too rigid overall. I've interviewed a handful of people for our team in the last year - all had the right background and technical expertise, but since they didn't have an example from their career which matches up with the LP questions they were asked, they were knocked back. All of them would have been "bar raising" from a technical standpoint, but since they were not demonstrably already "Amazonian" in their mindset, we couldn't hire them.

Re: AWS's us-east-1 region is experiencing issues

#150
post #137

Earlier quoted context omitted.

Wow, that just seems completely performative. Is "SA" systems admin here?

Solutions Architect, jack of all trades that is a sort of customer consultant and creator of solutions architecture (duh) for customers.

Customers and partners, the latter basically taking our ideas and selling them as a product - the flow of IP towards Amazon isn't always as clear-cut as people believe :-)

Broadly though it is a pre-sales role to help people get started, followed by ongoing guidance as the customer iterates (this is the part which often turns into free support).

Post reply on HN