Earlier quoted context omitted.
It’s standard. Career ladder [1] sets expectation for each level. Performance is measured against those expectations. Outages don’t negatively impact a single engineer. The key difference is the perspective. If reliability is bad that’s an organizational problem and blaming or punishing one engineer won’t fix that. [1] An example ladder from Patreon: https://levels.patreon.com/
> The key difference The key difference between what and what?
AWS us-east-1 outage
981–990 of 1001 posts
Re: AWS us-east-1 outage
#982Re: AWS us-east-1 outage
#983Earlier quoted context omitted.
They aren't meant to be, but shitty teams are shitty. You can also create a COE and assign it to another team. When I was at AWS, I had a few COEs assigned to me by disgruntled teams just trying to make me suffer and I told them to pound sand. For my own team, I wrote COEs quite often and found it to be a really great process for surfacing systemic issues with our management chain and making real improvements, but it…
At some point the number of people who were on shitty teams becomes an indictment on the wider culture at Amazon.
Re: AWS us-east-1 outage
#984Earlier quoted context omitted.
It's popular to upvote this during outages, because it fits a narrative. The truth (as always) is more complex: * No, this isn't the broad culture. It's not even a blip. These are EXCEPTIONAL circumstances by extremely bad teams that - if and when found out - would be intervened dramatically. * The broad culture is blameless post-mortems. Not whose fault is it. But what was the problem and how to fix it. And one of t…
100 BEZOBUCKS™ have been deposited to your account for this post.
Edit: also, could you please stop posting unsubstantive comments generally? You've unfortunately been doing that repeatedly, and we're trying for something else here.
Re: AWS us-east-1 outage
#985Earlier quoted context omitted.
> Should I be penalized if an upstream dependency, owned by another team, fails? Yes > Did I lack due diligence in choosing to accept the risk that the other team couldn't deliver? Yes
Where does this mindset end? Do I lack due diligence by choosing to accept that the cpu microcode on the system I’m deploying to works correctly?
Re: AWS us-east-1 outage
#986Earlier quoted context omitted.
Not in healthy organizations, they don't.
Once you reach a certain size there are surprisingly few healthy organization, most of them turn into externalization engines with 4 beats per year.
Re: AWS us-east-1 outage
#987Earlier quoted context omitted.
Pedantic, -1.
I suggest you review this before commenting again: https://news.ycombinator.com/newsguidelines.html
You mean like that rule, pedant? It's not name calling if it's an accurate representation of one's behavior.
ped·ant /ˈpednt/: noun a person who is excessively concerned with minor details and rules or with displaying academic learning.
Re: AWS us-east-1 outage
#988Earlier quoted context omitted.
Agreed - in my line of work regulators want everything in the country we operate from but of course CloudFront has to be different.
Wouldn't using a global CDN for everything be off the table to begin with, in that case?
Re: AWS us-east-1 outage
#989Earlier quoted context omitted.
> This issue is affecting the global console landing page, which is also hosted in US-EAST-1 Even this little tidbit is a bit of a wtf for me. Why do they consider it ok to have anything hosted in a single region? At a different (unnamed) FAANG, we considered it unacceptable to have anything depend on a single region. Even the dinky little volunteer-run thing which ran https://internal.site.example/~someEngineer was…
MAANG* How long before Meta takes over for Facebook?
Especially when you are getting burned by an outage.