Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

621–630 of 1001 posts

Re: AWS us-east-1 outage

#621
post #509

Earlier quoted context omitted.

It's popular to upvote this during outages, because it fits a narrative. The truth (as always) is more complex: * No, this isn't the broad culture. It's not even a blip. These are EXCEPTIONAL circumstances by extremely bad teams that - if and when found out - would be intervened dramatically. * The broad culture is blameless post-mortems. Not whose fault is it. But what was the problem and how to fix it. And one of t…

Can’t comment on most of your post but I know a lot of Amazon engineers who think of the CoE process (Correction of Error, what other companies would call a postmortem) as punitive

They aren't meant to be, but shitty teams are shitty. You can also create a COE and assign it to another team. When I was at AWS, I had a few COEs assigned to me by disgruntled teams just trying to make me suffer and I told them to pound sand. For my own team, I wrote COEs quite often and found it to be a really great process for surfacing systemic issues with our management chain and making real improvements, but it needs to be used correctly.

Re: AWS us-east-1 outage

#622
post #617

Earlier quoted context omitted.

Man, between costs and shenanigans like this, why don't more companies self-host?

Opex > Capex. If companies thought about long term, yes they might consider it. But unless the cloud providers fuck up really badly, they're ok to take the heat occasionally and tolerate a bit of nonsense.

You can lease equipment you know…

Re: AWS us-east-1 outage

#624
post #328

Earlier quoted context omitted.

I've worked for Amazon for 4 years, including stints at AWS, and even in my current role my team is involved in LSE's. I've never seen this behavior, the general culture has been find the problem, fix it, and then do root cause analysis to avoid it again. Jeff himself has said many times in All Hands and in public "Amazon is the best place to fail". Mainly because things will break, it's not that they break that's in…

I guess the question is why can't you (AWS) fix the problem of the status page not reflecting an outage? Maybe acceptable if the console has a hiccup, but when www.amazon.com isn't working right, there should be some yellow and red dots out there. With the size of your customer base there were man years spent confirming the outage after checking the status.

Because there's a VP approval step for updating the status page and no repercussions for VPs who don't approve updates in a timely manner. Updating the status page is fully automated on both sides of VP approval. If the status page doesn't update, it's because a VP wouldn't do it.

Re: AWS us-east-1 outage

#625

Earlier quoted context omitted.

If you're not multi-cloud in 2021 and are expecting 5-9's, I feel bad for you.

I imagine there are very few businesses where the extra cost of going multi-cloud is smaller than the cost of being down during AWS outages.

Also, going multi-cloud will introduce more complexity which leads to more errors and more downtime. I'd rather sit this outage out than deal with daily risk of downtime because I'm infrastructure is too smart for its own good.

Re: AWS us-east-1 outage

#626

Earlier quoted context omitted.

If you're not multi-cloud in 2021 and are expecting 5-9's, I feel bad for you.

If you're not multi-region, I feel bad for you. If your company is shoehorning you into using multiple clouds and learning a dozen products, IAM and CICD dialects simultaneously because "being cloud dependent is bad", I feel bad for you. Doing one cloud correctly from a current DevSecOps perspective is a multi-year ask. I estimate it takes about 25 people working full time on managing and securing infrastructure per…

This.

Re: AWS us-east-1 outage

#627
post #375

Earlier quoted context omitted.

It's because none of these companies are held responsible for missing their actual SLAs, as opposed to their self-reported SLA compliance. So unless regulation gets implemented that says otherwise, there's zero incentive for any company to maintain an accurate status page.

How did you find a way to bring regulations into this? There are monitoring services you can pay for to keep an eye on your SLAs and your vendors'. If not happy with the results switch.

Technically, there are already regulations. SLA lies are fraud.

But I'm leery of any business who's so dishonest they fear any outside oversight that brings repercussions for said dishonesty.

"If not happy, switch" is silly - it's not the customer's problem. And if you're a large customer and have invested heavily in getting staff trained on AWS, you can't just move.

Re: AWS us-east-1 outage

#628
post #618

this is exact kinda over centralisation issues I was talking about. I'm one of first developers using AWS EC2, sure back when scaling is hard for small dev shops. In now day and age any one who is technically inclined, can figure out using the new technologies. Why even use AWS. Get something like Hetzner, Linodes please!

Some manager: But it does not web scale!

the average HN poster: I run things on a box under my desk and you cannot teach me why thats bad!!!

Re: AWS us-east-1 outage

#629
imdb seems down too and returning 503. Is it related? Here is the output. Kind of funny.

D'oh!

Error 503

We're sorry, something went wrong.

Please try again...wait...wait...yep, try reload/refresh now.

But if you are seeing this again, please report it here.

Please explain which page you were at and where on it that you clicked

Thank you!

Post reply on HN