Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

941–950 of 1001 posts

Re: AWS us-east-1 outage

#942

Haha my developer called me in panic telling that he crashed Amazon - was doing some load tests with Lambda

Postmortem: unbounded auto-scaling of lambda combined with oversight on internal rate limits caused unforseen internal ddos.

He created a lambda function that spawned more lambda functions and the rest is history

Re: AWS us-east-1 outage

#943
post #310

Earlier quoted context omitted.

I don't think engineers can believe in no-blame analysis if they know it'll harm career growth. I can't unilaterally promote John Doe, I have to convince other leaders that John would do well the next level up. And in those discussions, they could bring up "but John has caused 3 incidents this year", and honestly, maybe they'd be right.

You know why SO-teams, firefighters and military pilots are so successful? -You don't hide anything -Errors will be made -After training/mission everyone talks about the errors (or potential ones) and how to prevent them -You don't make the same error twice Being afraid to make errors and learn from them creates a culture of hiding, a culture of denial and especially being afraid to take responsibility.

Yes. AAR process in the army was good at this up to the field grade level, but got hairy on G/J level staffs. I preferred being S-6 to G-6 for that reason.

Re: AWS us-east-1 outage

#944

Earlier quoted context omitted.

> This issue is affecting the global console landing page, which is also hosted in US-EAST-1 Even this little tidbit is a bit of a wtf for me. Why do they consider it ok to have anything hosted in a single region? At a different (unnamed) FAANG, we considered it unacceptable to have anything depend on a single region. Even the dinky little volunteer-run thing which ran https://internal.site.example/~someEngineer was…

MAANG* How long before Meta takes over for Facebook?

MANGA

Re: AWS us-east-1 outage

#945
post #700

The fun thing about these types of outages are seeing all of the people that depend upon these services with no graceful fallback. My roomba app will not even launch because of the AWS outage. I understand that the app gets "updates" from the cloud. In this case "updates" is usually promotional crap, but whatevs. However, for this to prevent the app launching in a manner that I can control my local device is total BS…

there's a reason it is called the Internet of things, and not the "local network of things". Even if the latter is probably what most customers would prefer.

The internet was originally designed to still work if there was a nuclear strike on a node. Now we can’t even cope with us-west-1 down.

Re: AWS us-east-1 outage

#946
post #828
post #779

Earlier quoted context omitted.

Apple created their own silicon. Fedex uses its own pilots. The USPS uses it's own cars. If you're a company relying upon AWS for your business, is it okay if you're down for a day, or two while you wait for AWS to resolve it's issue?

> Apple created their own silicon. Apple designs the M1. But TSMC (and possibly Samsung) actually manufacture the chips.

I'm pretty sure that's a difference without a distinction.

Re: AWS us-east-1 outage

#947

Earlier quoted context omitted.

So on what basis is someone's performance reviewed, if such performance is omitted?

The entire point of blameless postmortems is acknowledging that the mere existence of an outage does not inherently reflect on the performance of the people involved. This allows you to instead focus on building resilient systems that avoid the possibility of accidental outages in the first place.

I know. That's not what I'm asking about, if you might read my question.

Re: AWS us-east-1 outage

#948
post #888

Earlier quoted context omitted.

So on what basis is someone's performance reviewed, if such performance is omitted?

I'll play devil's advocate here and say that sometimes these incidents deserve praise because they uncovered an issue that was otherwise unknown previously. Also if the incident had a large negative impact then it shows to leadership how critical normal operation of that service is. Even if you were the cause of the issue, the fact that you fixed it and kept the critical service operating the rest of the time, is wor…

I know; that's not what I'm asking about. I'm talking about a different issue.

Re: AWS us-east-1 outage

#949
post #903
post #872

Earlier quoted context omitted.

> I would imagine that the US IC (CIA/NSA) already does some free consulting for these giant companies This comment is how I know you don't work in the public sector. Those agencies' infrastructures are essentially run by contractors with a few GS personnel making bad decisions every chance they get and a few DoD personnel acting like their rank can fix technical problems.

I'm not talking about running infrastructure, I'm talking about working with HR and making sure that the people they've hired as sysadmins aren't meeting with FSB or PLA agents in a local park on weekends and accepting suitcases of USD cash to accidentally `rm -rf` all of the Zookeeper/etcd nodes at once on a Monday morning.

I doubt it. We’re shit at critical infrastructure defense and the military cares mostly about its own networks. And industry really doesn’t want to cooperate. I was in cyber and military IT. Can’t speak for the IC, but I really doubt it.

Re: AWS us-east-1 outage

#950
post #863
post #839

Earlier quoted context omitted.

>Facepalm in the form of $100Ms in service credits. Part of me wonders how much they're actually going to pay out, given that their own status page has only indicated five services with moderate ("Increased API Error Rates") disruptions in service.

Utter lies on that page. Multiple services listed as green aren't working for me or my team.

https://stop.lying.cloud
Post reply on HN