Live data from Hacker News

Tell HN: AWS appears to be down again

news.ycombinator.com

321–330 of 646 posts

Re: Tell HN: AWS appears to be down again

#321
post #296

Earlier quoted context omitted.

forgive me repeating myself: AWS Zones are not truly independent of each other. Global services such as route53, Cognito, the default cloud console and Cloudfront are managed out of US-East-1. If us-east-1 is unavailable, as is commonly the case, and you depend on those systems, you are also down. it does not matter if you're in timbuktu-1, you are dead in the water. it is a myth that amazon availability zones are tr…

Of course that depends on what services you use and yes, even then there is some remaining correlation just because it is the same host. > are not truly independent of each other Indeed. They are even on the same planet! > please stop blaming the victim Excuse me?

>> are not truly independent of each other

> Indeed. They are even on the same planet!

Clever bastard, aren't you.

>> please stop blaming the victim

> Excuse me?

"If you're affected by us-east-1 outages then you're not hosting in other regions and you're doing it wrong".

Except: You can be affected by this outage if you did everything right. You're putting blame on people being down for not being hosted in different regions when it would not help them. You've effectively shifted blame away from Amazon and onto the person who cannot control their uptime by doing what you said.

Re: Tell HN: AWS appears to be down again

#322

Earlier quoted context omitted.

> Can you make your on prem infrastructure go down less than Amazon's? Obviously depends on what you need, but for a small to medium web app that needs a load-balancer, a few app servers, a database and a cache, yes absolutely - all of these have been solved problems for over a decade and aren't rocket science to install & maintain. > Is it worth it? I'd argue that the "worth" would be less about immunity to occasion…

Human capital side would disagree with that I think. You're assuming the organization which owns this small/medium web app has the personnel already on staff to handle such a thing. If you're outsourcing that, you'd likely have to pay a boatload just for someone to be available for help, let alone the actual tasks themselves. Like you said, if you're on-prem and something goes down, you can do something. But you've g…

> Human capital side would disagree with that I think

I hear this argument a lot, but every startup I've been involved with had a full-time DevOps engineer wrangling Terraform & YAML files - that same engineer can be assigned to manage the bare-metal infrastructure.

Re: Tell HN: AWS appears to be down again

#323
post #60

So, how many execs are going to push to move to self-managed hosting in the new year? Packaging a way to migrate off AWS could be a unicorn idea.

Would need one hell of a compressional algorithm to keep the data exfiltration costs down.

Pied Piper

Re: Tell HN: AWS appears to be down again

#324
post #154

The prevailing wisdom throughout the last couple of years was: “ditch your on-prem infrastructure and migrate to a major cloud provider” And its starting to seem like it could be something like: “ditch your on-prem infrastructure and spin up your own managed cloud” This is probably untenable for larger orgs where convenience gets the blank check treatment, but for smaller operations that can’t realize that value at s…

“Hybrid and multi cloud” is the future. In other words, give us more fucking money.

Re: Tell HN: AWS appears to be down again

#325

Earlier quoted context omitted.

> I've only had a DC fail once when the engineer was performing work on the power circuitry for the DC and thought he was taking down one, but was in fact the wrong one and took both power circuits down at the same time. This is all local scale. Your setup would not survive a data center scale power outage. At scale power outages are datacenter scale. Data centers lose supply lines. They lose transformers. Sometimes…

"It is cheaper to design a system that must be up which accounts for a data center being totally down and a portion of the system being totally unavailable than to add more datacenter mitigations." Citation needed - the same issue with testing, data races and expensive bandwidth come up.

At high energy the lead time for the components is measured not in days but in years.

Re: Tell HN: AWS appears to be down again

#326

Earlier quoted context omitted.

> Can you make your on prem infrastructure go down less than Amazon's? Obviously depends on what you need, but for a small to medium web app that needs a load-balancer, a few app servers, a database and a cache, yes absolutely - all of these have been solved problems for over a decade and aren't rocket science to install & maintain. > Is it worth it? I'd argue that the "worth" would be less about immunity to occasion…

Human capital side would disagree with that I think. You're assuming the organization which owns this small/medium web app has the personnel already on staff to handle such a thing. If you're outsourcing that, you'd likely have to pay a boatload just for someone to be available for help, let alone the actual tasks themselves. Like you said, if you're on-prem and something goes down, you can do something. But you've g…

You still need to pay someone to manage AWS infrastructure. It’s possible to save money using AWS, but things often get more expensive.

Re: Tell HN: AWS appears to be down again

#327
post #74

Earlier quoted context omitted.

> I'm not sure if we should say "AWS is down" if only us-east-1 is down. The thing is, us-east-1 represents the whole AWS for the majority of us.

Can you expand on that? What feature do you use in east 1 that isn’t everywhere else that it’s your whole implementation?

> Can you expand on that? What feature do you use in east 1 that isn’t everywhere else that it’s your whole implementation?

Your question reads as a strawman. It matters nothing if EC2 is also available in Mumbai or Hong Kong if by default the whole world deploys everything and anything to us-east-1, and us-east-1 alone.

https://www.reddit.com/r/aws/comments/nztxa5/why_useast1_reg...

Re: Tell HN: AWS appears to be down again

#328

Me: Hesitation at last job moving absolutely everything (including backups) to AWS because if it goes down it's a problem I'm a firm believer in some kind of physical/easily accessible backup. Coworkers: "You're an f'n idiot. Amazon and Facebook don't go down, you're holding us back!" Me: leaves cause that treatment was the final straw Amazon and Facebook both go down within a month of each other, and supposedly they…

Seems like multi-cloud solution might be the way to go.

All while making sure that these cloud solutions are not inter-dependent and that there are redundant paths to access these services.

Re: Tell HN: AWS appears to be down again

#330

Earlier quoted context omitted.

Think about it this way: 1) Can you make your on prem infrastructure go down less than Amazon's? 2) Is it worth it? In my experience most people grossly underestimate how expensive it is to create reliable infrastructure and at the same time overestimate how important it is for their services to run uninterrupted. -- EDIT: I am not arguing you shouldn't build your more reliable infrastructure. AWS is just a point on…

> Can you make your on prem infrastructure go down less than Amazon's? Obviously depends on what you need, but for a small to medium web app that needs a load-balancer, a few app servers, a database and a cache, yes absolutely - all of these have been solved problems for over a decade and aren't rocket science to install & maintain. > Is it worth it? I'd argue that the "worth" would be less about immunity to occasion…

I have run high availability (HA) systems in prem and your statement vastly understates the difficulty and expense.

You need multiple physical links in running to different ISPs because builders working on properties further down the street could accidentally cut through your fibre. Or the ISP themselves could suffer an outage.

You need a back up generator and to be a short distance away from a petrol station so you can refuel quickly and regularly when suffering from longer durations of power outages. You absolutely do not want to run out of diesel!

You need redundancy of every piece of hardware AND you need to test that failover works as expected because the last thing you need is a core switch to fail and traffic not to route over secondary core switch like expected.

You need your multiple air con units and them to be powered off different mains inputs so if the electrics fail on one unit it doesn’t take out the others. I guarantee you that if the air cons will fail, it will be on the hottest day of the year a month amount of portable units will stop your servers from overheating.

You need beefy UPS with multiple batteries. Ideally multiple UPSs with each UPS powering a different rail on your racks so that if one UPS fails your hardware is still powered from the other rail. And you need to regularly check the battery status and loads on the UPS. Remember that the back up generator takes a second or two to kick in so you need something to keep the power to the servers and networking hardware to be uninterrupted. And since all your hardware is powered via the UPS, if that dies you still lose power even if the building is powered.

And you then need to duplicate all of the above in second location just in case the first location still goes down.

By the way, all of the possible failure points I’ve raised above HAVE failed on me when managing HA on prem.

The reason people move to the cloud for HA is because rolling your own is like rolling your own encryption: it’s hard, error prone, expensive, and even when you have the right people on the team there’s still a good chance you’ll fuck it up. AWS, for all its faults, does make this side of the job easier.

Post reply on HN