Live data from Hacker News

AWS us-east-1 outage

status.aws.amazon.com

611–620 of 1001 posts

Re: AWS us-east-1 outage

#611

Earlier quoted context omitted.

Seems "the cloud" had a major outage less than a month ago, my laptop has a higher uptime. $ 16:04 up 46 days, 7:02, 9 users, load averages: 3.68 3.56 3.18 US East 1 was down just over a year ago https://www.theregister.com/2020/11/25/aws_down/ Meanwhile I moved one of my two internal DNS servers to a second site on 11 Nov 2020, and it's been up since then. One of my monitoring machines has been filling, rotating and…

so does my toaster, oven and microwave. so what? they get used a few times a day, but my production level equipment serves millions in an hour.

My lightswitch is used twice a day, yet it works every time. In the old days it would occasionally break (bulb goes), I would be empowered to fix it myself (change the bulb).

In the cloud you're at the mercy of someone who doesn't even know you exist to fix it, without the protections that say an electric company has with supplying domestic users.

This thread has people unable to turn their lights on[0], it's hilarious how people tie their stuff to dependencies that aren't needed, with a history of constant failure.

If you want to host millions of people, then presumably your infrastructure can cope with the loss of a single AZ (and ideally the loss of Amazon as a whole). The vast majority of people will be far better off without their critical infrastructure going down in the middle of the day in the busiest sales season going.

[0] https://news.ycombinator.com/item?id=29475499

Re: AWS us-east-1 outage

#612

Earlier quoted context omitted.

You and many others here may be conflating two concepts which are actually quite separate. Taking blame is a purely punitive action and solves nothing. Taking responsibility means it's your job to correct the problem. I find that the more "political" the culture in the organization is, the more likely everyone is to search for a scapegoat to protect their own image when a mistake happens. The higher you go up in the…

Every argument I have on the internet is between prescriptive and descriptive language. People tend to believe that if you can describe a problem that means you can prescribe a solution. Often times, the only way to survive is to make it clear that the first thing you are doing is describing the problem. After you do that, and it's clear that's all you are doing, then you follow up with a prescriptive description whe…

My comment was made from the relatively simpler entrepreneurial perspective, not the corporate one. Corp ownership rests with people in the C-suite who are social/political lawyer types, not technical people. They delegate responsibility but not authority, because they can hire people, even smart people, to work under those conditions. This is an error mode where "blame" flows from those who control the money to those who control the technology. Luckily, not all money is stupid so some corps (and some parts of corps) manage to function even in the presence of risk and innovation failures. I mean the whole industry is effectively a distributed R&D budget that may or may not yield fruit. I suppose this is the market figuring out whether iterated R&D makes sense or not. (Based on history, I'd say it makes a lot of sense.)

Re: AWS us-east-1 outage

#613
this is exact kinda over centralisation issues I was talking about. I'm one of first developers using AWS EC2, sure back when scaling is hard for small dev shops. In now day and age any one who is technically inclined, can figure out using the new technologies. Why even use AWS. Get something like Hetzner, Linodes please!

Re: AWS us-east-1 outage

#614

Earlier quoted context omitted.

if I were a black hat I would absolutely love GitHub and all the various language-specific package systems out there. giving me sooooo many ways to sneak arbitrary tailored malicious code into millions of installs around the world 24x7. sure, some of my attempts might get caught, or not but not lead to a valuable outcome for me. but that percentage that does? can make it worth it. its about scale and a massive parall…

Any yet oddly enough the Earth continues to spin and the internet continues to work. I think the system we have now is necessarily the system that must exist ( in this particular case, not in all cases ). Something more centralized is destined to fail. And, while the open source nature of software introduces vulnerabilities it also fixes them.

> And, while the open source nature of software introduces vulnerabilities it also fixes them.

dat gap tho... which was my point. smart black hats will be exploiting this gap, at scale. and the strategy will work because the majority of folks seem to be either lazy, ignorant or simply hurried for time.

and btw your 1st sentence was rude. constructive feedback for the future

Re: AWS us-east-1 outage

#615

Earlier quoted context omitted.

Why would I want to triple my capacity? Most people don't need to scale to a billion users overnight.

Many B2B-type applications have a lot of usage during the workday and minimal usage outside of it. No reason to keep all that capacity running 24/7 when you only need most of it for ~8 hours per weekday. The cloud is perfect for that use case.

Is it really? How much does that scaling actually cost?

And what's a workday anyway, surely you operate globally?

Re: AWS us-east-1 outage

#616

Earlier quoted context omitted.

Why would I want to triple my capacity? Most people don't need to scale to a billion users overnight.

Many B2B-type applications have a lot of usage during the workday and minimal usage outside of it. No reason to keep all that capacity running 24/7 when you only need most of it for ~8 hours per weekday. The cloud is perfect for that use case.

[deleted]

Re: AWS us-east-1 outage

#617

Earlier quoted context omitted.

One time gcp argued that since they did return 404s on gcs for a few hours that wasn’t an uptime/latency sla violation so we were not entitled to refund (tho they refunded us anyway)

Man, between costs and shenanigans like this, why don't more companies self-host?

Opex > Capex. If companies thought about long term, yes they might consider it. But unless the cloud providers fuck up really badly, they're ok to take the heat occasionally and tolerate a bit of nonsense.

Re: AWS us-east-1 outage

#618

this is exact kinda over centralisation issues I was talking about. I'm one of first developers using AWS EC2, sure back when scaling is hard for small dev shops. In now day and age any one who is technically inclined, can figure out using the new technologies. Why even use AWS. Get something like Hetzner, Linodes please!

Some manager: But it does not web scale!

Re: AWS us-east-1 outage

#620

Earlier quoted context omitted.

It's popular to upvote this during outages, because it fits a narrative. The truth (as always) is more complex: * No, this isn't the broad culture. It's not even a blip. These are EXCEPTIONAL circumstances by extremely bad teams that - if and when found out - would be intervened dramatically. * The broad culture is blameless post-mortems. Not whose fault is it. But what was the problem and how to fix it. And one of t…

If your statement is true, then why is the AWS status page widely considered useless, and everyone congregates on HN and/or Twitter to actually know what's broken on AWS during an outage?

> Yes, VP approval is needed to make any updates on the status dashboard. But that's not as hard as it may seem. AWS executives are extremely operation-obsessed, and when there is an outage of any size are engaged with their service teams immediately.

My experience generally aligns with amzn-throw, but this right here is why. There's a manual step here and there's always drama surrounding it. The process to update the status page is fully automated on both sides of this step, if you removed VP approval, the page would update immediately. So if the page doesn't update, it is always a VP dragging their feet. Even worse is that lags in this step were never discussed in the postmortem reviews that I was a part of.

Post reply on HN