Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

861–870 of 1001 posts

Re: AWS us-east-1 outage

#861

Earlier quoted context omitted.

When I was at Google I didn't have a lot of exposure to the public infra side. However I do remember back in 2008 when a colleague was working on routing side of YouTube, he made a change that cost millions of dollars in mere hours before noticing and reverting it. He mentioned this to the larger team which gave applause during a tech talk. I cannot possibly generalize the culture differences between Amazon and Googl…

While I support that, how are the people involved evaluated?

I was not informed of his performance reviews. However, given the reception, his work in general, and the attitudes of the team, I cannot imagine this even came up. More likely the ability to improve routing to actually make YouTube cheaper in the end was I'm sure the ultimate positive result.

This was also towards the end of the golden age of Google, when the percentage of top talent was a lot higher.

Re: AWS us-east-1 outage

#862
post #245

Looks like they've acknowledged it on the status page now. https://status.aws.amazon.com/ > 8:22 AM PST We are investigating increased error rates for the AWS Management Console. > 8:26 AM PST We are experiencing API and console issues in the US-EAST-1 Region. We have identified root cause and we are actively working towards recovery. This issue is affecting the global console landing page, which is also hosted in US…

As a user of Sagemaker in us-east-1, I deeply fucking resent AWS claiming the service is normal. I have extremely sensitive data, so Sagemaker notebooks and certain studio tools make sense for me. Or DID. After this I'm going back to my previous formula of EC2 and hosting my own GPU boxes.

Sagemaker is not working, I can't get to my work (notebook instance is frozen upon launch, with zero way to stop it or restart it) and Sagemaker Studio is also broken right now.

The length of this outage has blown my mind.

Re: AWS us-east-1 outage

#863
post #839

Earlier quoted context omitted.

Every damn Well-Architected Framework includes multi-AZ if not multi-region redundancy, and yet the single access point for their millions of customers is single-region. Facepalm in the form of $100Ms in service credits.

>Facepalm in the form of $100Ms in service credits. Part of me wonders how much they're actually going to pay out, given that their own status page has only indicated five services with moderate ("Increased API Error Rates") disruptions in service.

Utter lies on that page. Multiple services listed as green aren't working for me or my team.

Re: AWS us-east-1 outage

#866
post #779

Earlier quoted context omitted.

Apple created their own silicon. Fedex uses its own pilots. The USPS uses it's own cars. If you're a company relying upon AWS for your business, is it okay if you're down for a day, or two while you wait for AWS to resolve it's issue?

It’s bloody annoying when all I want to do is vacuum the floor and Roomba says nope, “active AWS incident”.

When I bought my automated sprinkler system, I got one that would continue to work if the company or the cloud went belly up.

Re: AWS us-east-1 outage

#868

Earlier quoted context omitted.

I am finding that I have a very bimodal response to "He did it". When I write an RCA or just talk about near misses, I may give you enough details to figure out that Tom was the one who broke it, but I'm not going to say Tom on the record anywhere, with one extremely obvious exception. If I think Tom has a toxic combination of poor judgement, Dunning-Kruger syndrome, and a hint of narcissism (I'm not sure but I may b…

To me, the point of "blameless" PM is not to hide the identity of the person who was closest to the failure point. You can't understand what happened unless you know who did what, when. "Blameless" to me means you acknowledge that the ultimate problem isn't that someone made a mistake that caused an outage. The problem is that you had a system in place where someone could make a single mistake and cause an outage. If…

That's true if the direct cause is an actual mistake, which often is the case but not always.

It may also be that the cause is willful negligence, intentionally circumventing barriers for some personal reason.

And, of course, it may be that the cause is explicitly malicious (e.g. internal fraud, or the intent to sabotage someone) and at least part of the blame directly lies on the culprit, and not only on those who failed to notice and stop them.

Re: AWS us-east-1 outage

#869

Earlier quoted context omitted.

While I support that, how are the people involved evaluated?

I was not informed of his performance reviews. However, given the reception, his work in general, and the attitudes of the team, I cannot imagine this even came up. More likely the ability to improve routing to actually make YouTube cheaper in the end was I'm sure the ultimate positive result. This was also towards the end of the golden age of Google, when the percentage of top talent was a lot higher.

So on what basis is someone's performance reviewed, if such performance is omitted?

Re: AWS us-east-1 outage

#870
post #863
post #839

Earlier quoted context omitted.

>Facepalm in the form of $100Ms in service credits. Part of me wonders how much they're actually going to pay out, given that their own status page has only indicated five services with moderate ("Increased API Error Rates") disruptions in service.

Utter lies on that page. Multiple services listed as green aren't working for me or my team.

Suggesting that when the status page sends a status request and hears no response—it defaults to green—hear no evil and see no evil —> report no evil

Either way—overt lies or engineering incompetence—it’s disappointing!

Post reply on HN