Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

981–990 of 1001 posts

Re: AWS us-east-1 outage

#981

Earlier quoted context omitted.

It’s standard. Career ladder [1] sets expectation for each level. Performance is measured against those expectations. Outages don’t negatively impact a single engineer. The key difference is the perspective. If reliability is bad that’s an organizational problem and blaming or punishing one engineer won’t fix that. [1] An example ladder from Patreon: https://levels.patreon.com/

> The key difference The key difference between what and what?

Your approach and their approach. It sounded like you have a different perspective about who is responsible for an outage.

Re: AWS us-east-1 outage

#983

Earlier quoted context omitted.

They aren't meant to be, but shitty teams are shitty. You can also create a COE and assign it to another team. When I was at AWS, I had a few COEs assigned to me by disgruntled teams just trying to make me suffer and I told them to pound sand. For my own team, I wrote COEs quite often and found it to be a really great process for surfacing systemic issues with our management chain and making real improvements, but it…

At some point the number of people who were on shitty teams becomes an indictment on the wider culture at Amazon.

Absolutely! Anecdotally, out of all the teams I interacted with in seven years at AWS across multiple arms of the company, I saw only a handful of bad teams. But like online reviews, the unhappy people are typically the loudest. I'm happy they are though, it's always important to push to be better, but I don't believe that AWS is the hellish place to work that HN discourse would lead many to believe.

Re: AWS us-east-1 outage

#984

Earlier quoted context omitted.

It's popular to upvote this during outages, because it fits a narrative. The truth (as always) is more complex: * No, this isn't the broad culture. It's not even a blip. These are EXCEPTIONAL circumstances by extremely bad teams that - if and when found out - would be intervened dramatically. * The broad culture is blameless post-mortems. Not whose fault is it. But what was the problem and how to fix it. And one of t…

100 BEZOBUCKS™ have been deposited to your account for this post.

This post broke the site guidelines badly. If you'd please review https://news.ycombinator.com/newsguidelines.html and stick to the rules in the future, we'd be grateful.

Edit: also, could you please stop posting unsubstantive comments generally? You've unfortunately been doing that repeatedly, and we're trying for something else here.

Re: AWS us-east-1 outage

#985

Earlier quoted context omitted.

> Should I be penalized if an upstream dependency, owned by another team, fails? Yes > Did I lack due diligence in choosing to accept the risk that the other team couldn't deliver? Yes

Where does this mindset end? Do I lack due diligence by choosing to accept that the cpu microcode on the system I’m deploying to works correctly?

I think this is why we pay for support, with the expectation that if their product inadvertently causes losses for you they will work fast to fix it or cover the losses.

Re: AWS us-east-1 outage

#986
post #276

Earlier quoted context omitted.

Not in healthy organizations, they don't.

Once you reach a certain size there are surprisingly few healthy organization, most of them turn into externalization engines with 4 beats per year.

I love it when I share a mental model with someone in the wild.

Re: AWS us-east-1 outage

#987
post #951

Earlier quoted context omitted.

Pedantic, -1.

I suggest you review this before commenting again: https://news.ycombinator.com/newsguidelines.html

> Please respond to the strongest plausible interpretation of what someone says, not a weaker one that's easier to criticize. Assume good faith.

You mean like that rule, pedant? It's not name calling if it's an accurate representation of one's behavior.

ped·ant /ˈpednt/: noun a person who is excessively concerned with minor details and rules or with displaying academic learning.

Re: AWS us-east-1 outage

#988

Earlier quoted context omitted.

Agreed - in my line of work regulators want everything in the country we operate from but of course CloudFront has to be different.

Wouldn't using a global CDN for everything be off the table to begin with, in that case?

Apparently it's okay for static data (like a website hosted in S3 behind CloudFront) but seeing non-Australian items in AWS billing and overviews always makes us look twice.

Re: AWS us-east-1 outage

#989

Earlier quoted context omitted.

> This issue is affecting the global console landing page, which is also hosted in US-EAST-1 Even this little tidbit is a bit of a wtf for me. Why do they consider it ok to have anything hosted in a single region? At a different (unnamed) FAANG, we considered it unacceptable to have anything depend on a single region. Even the dinky little volunteer-run thing which ran https://internal.site.example/~someEngineer was…

MAANG* How long before Meta takes over for Facebook?

I like MAGMA (Meta, Amazon, Google, Microsoft, Apple).

Especially when you are getting burned by an outage.

Re: AWS us-east-1 outage

#990
post #860

Earlier quoted context omitted.

One region? I forgot how to count that low

It's like three regions - when two of them explode. Two is one & one is none.

the obvious solution is to put all internet in one region so that when that one explodes nobody notices your little service
Post reply on HN