Live data from Hacker News

AWS us-east-1 down

news.ycombinator.com

281–290 of 331 posts

Re: AWS us-east-1 down

#281
post #268

Earlier quoted context omitted.

Are you just guessing it's related or is there a reason to be sure? Either way, does seem sjll6 to have no manual fallback 'check-in' (or policy-level ability to bypass the need for it) mechanism.

Most retail stores simply stop if "the system" is down.

I've known 'just take your shopping and go' and 'let me make a note of your details', but never 'put your bags down and bugger off but please do comw again'.

And why would ordered and paid for shopping be so inaccessible? It's not going to be locked away in a vending machine type thing right, you just need to show up with ID/log-in and collect it. Or so you should.

Re: AWS us-east-1 down

#283
post #281

Earlier quoted context omitted.

Most retail stores simply stop if "the system" is down.

I've known 'just take your shopping and go' and 'let me make a note of your details', but never 'put your bags down and bugger off but please do comw again'. And why would ordered and paid for shopping be so inaccessible? It's not going to be locked away in a vending machine type thing right, you just need to show up with ID/log-in and collect it. Or so you should.

I’ve only known “sorry, the till is down” and line ups as far as the eye can see.

It's ridiculous that everything stops, but I do understand.

The pickup is even more ridiculous. Like you say, it's already through the system. They would just need to mark it as delivered later.

Re: AWS us-east-1 down

#284

Why is it always us-east-1 though? I have always stayed away from that region because it seems significantly less reliable than other regions.

I actually just wrote about this very thing. It's not just that it SEEMS less reliable, it absolutely is: https://statusgator.com/blog/is-north-virginia-aws-region-th...

I don't think this article has any value. Are you only counting region wide outages? US east is probably 10x the size of any other region with more AZ's than any other region.

Re: AWS us-east-1 down

#285
post #85

Earlier quoted context omitted.

I mostly struggle with Irish Standard Time (used for DST in Ireland) and Indian Standard Time which have the same acronym. :( Thankfully, I learnt a long time ago to use ISO 8601 and UTC for dates and times. I still revert to PST/PDT if my audience is primarily left coast based.

And I can't say it's ever actually caused a problem, but something about Indian Standard Time being a half-hour offset from UTC has always bothered me so much... But now we're fully off-topic.

Oh, hold on to something, while I tell you about Chatham Islands.

Re: AWS us-east-1 down

#286

Mysterious lack of "AWS is bad for the internet because it is so centralized" dialog up in here. edit: for those that would downvote: HN _just_ yesterday: https://news.ycombinator.com/item?id=36295352 https://news.ycombinator.com/item?id=36295305

Ok fine. Running your own datacenter in 2023 is incredibly risky. There's the upfront server cost and the ongoing maintenance cost. There's patches and staffing and disaster planning and all the other things that goes into it. Plus there's the cyberinsurance and protections and security components too. Do you really think other (smaller) orgs can do a better job at hosting a datacenter than Amazon / Google / Microsof…

> Running your own datacenter in 2023 is incredibly risky.

There are middle grounds.

But let's be honest: 99% of companies have never done the napkin math, because nobody ever got fired for choosing IBM^W AWS.

We joked about this in my company: we had a variable-load thing that we used autoscaling in the cloud for, but it had a baseline load that purchasing a real machine might have made a lot of sense for. The napkin math probably checked out. We never suggested it more than jokingly, though, because even when we suggested it jokingly, we got shut down: "You don't understand the cost of that." No, actually, we jokingly did enough math that we do understand, better than the people criticizing us did. We never did it.

Whenever the "own it" argument comes up, eveybody is real quick to hop on the "but maintenance cost" train. But as I perceive it, those who believe in the cloud budget exactly $0 for maintenance of managed cloud resources. As someone who's only done cloud, that number is unadulterated bullshit: the number of hours I've had to spend chasing cloud vendors to do the job that we're paying for is just silently flying under the budget radar. In the minds of the finance books, I'm 100% SWE, but in reality, I'm 75% SWE, and 25% support ticket monkey.

At least with a real machine, it'd be interesting, and I'd have some agency to actually solve the problem. As a support ticket monkey, I'm utterly powerless. I'm tired of having to beg.

That's not to say I'd move everything off cloud; I actually think the vast majority of what we do is well-suited for cloud, mostly because upper management can't make up their mind about product direction enough to be able to say "yes, we can purchase this and we'll use it." But those nuggets of stability do happen from time to time.

> disaster planning

"Disaster planning" is something every org wants, because they're trying to tick the box with the regulator. But the requirements that get passed down border on absurd: "what if a meteor hit AWS and they were never able to recover from it?" … we're literally never going to plan for that, because the $ needed for that level of eng. work is not going to happen. A sane scenario would be "can we handle an AZ outage?" (or, let's start there, and maybe, maybe if we can get that down pat, then we can graduate to regional outages.)

> cyberinsurance

… you don't get out of this via being in the cloud, if you need it. (I wish we did, because ours pushes some utter inane requirements.) I can mismanage a machine in a DC just as easily as I can mismanage a VM in the cloud.

> Organizations pay attention to dollars

No they don't. This oft-repeated mantra is nonsense. Finance dept. get an invoice that has a total; even were they to have access to the finer billing information, they're not technical, and cannot understand it. I've yet to be at a company that's dedicated sufficient resources towards infra eng such that we could do the legwork necessary to present a sane organizational view of what cloud infra dollars go to what high-level objectives or teams. The resource tagging isn't there, and even if it were, some things cannot be tagged, and you still have to aggregate bills from a dozen different vendors, and then figure out what weights to apply to shared resources across OUs. I'm on employer #4? and have yet to see anyone scratch the surface of that.

Which is why you see articles about cloud $ waste all the time.

What happens far more often in my life is someone from management descending with "why are we spending $X on Y?", where $X is usually an order of magnitude wrong, or Y is … something we're not even doing anymore? And then you have to go round the mulberry bush of "how did you arrive at that figure?" "okay so here's what those numbers mean" "here you're adding $/mo and $/yr and you can't do that"

> Do you really think other (smaller) orgs can do a better job at hosting a datacenter than Amazon / Google / Microsoft / Cloudflare?

Than Microsoft? Absolutely yes. The others, probably not.

> Yes, I get it. All the computer processing power in a handful of actor's hands is probably not the most fantastic thing.

The long-term end state of not investing money into R&D is that it is centralized into those who do, and you become beholden to them. You get what you pay for, here. It's not good, and I think there's discussion to be had around that, but my real problem is the cognitive dissonance that follows. If you want to centralize on one of the cloud duopoly, then you also need to acknowledge that your own eng cannot be held responsible for the cloud's reliability: they have no control over it.

Re: AWS us-east-1 down

#287
post #239

Earlier quoted context omitted.

Even if people _do_ care, there isn't much to do about it.

If people do care they could use other hosting providers such as Hetzner or OVH, no?

Why did you come to the conclusion that Hetzner or OVH is more reliable. At least their SLA credits doesn't say that.

Re: AWS us-east-1 down

#288

Why does everyone keep deploying their products to this one region when it always seems like the one that fails? We don't use big cloud were I work, so maybe I'm missing something. Does East-1 offer something other don't?

Anecdotal with n=2 sample, but GPU availability seems better in us-east-1.

Re: AWS us-east-1 down

#289

Earlier quoted context omitted.

There's a lot of software, iirc even Amazon's own dashboards, that simply defaults to us-east-1.

That's my favorite part of Amazon's console. That miniature heart attack you have when you ask "WHERE ARE ALL MY LAMBDAS AND DYNAMO INSTANCES?" Then you realize that they just switched you back to us-east-1 for some reason and a wave of familiar relief washes over you.

Out of curiosity have you been building with the serverless framework? I'm curious what drives people to use dynamodb in particular

Re: AWS us-east-1 down

#290

For those wondering: Currently PDT is 7 hours behind UTC. AWS can do so many things, reporting critical outage updates in UTC is not one of those things.

When I was with AWS I advocate for ISO8601 "Z" whenever I could or need to influence, say internal systems.

If all systems talk this we'd save tens of thousands of man hours. Just do the conversion for us mortals, or other necessities. Tech side of incidents is definitely "system", I'd argue more often than not consumers of AWS are also tech side with systems in UTCs so health dashboards should also be a UTC first system. Doubt this could get prioritized tho

Post reply on HN