Live data from Hacker News

Tell HN: AWS appears to be down again

news.ycombinator.com

631–640 of 646 posts

Re: Tell HN: AWS appears to be down again

#631
post #228

Earlier quoted context omitted.

> Not at this rate. Source? Has there ever been an industry wide survey that compares availability from "insert average colo/data center operations" with the cloud ones? And I'm not talking about "we have 12 SREs who are based in Cupertino and are all paid top dollar to support a colo"...I'm talking average .

I've worked at a few small companies over the years that had their own significant colos and/or data centers built on the cheap, and only a sysadmin or two to run them. Anecdotally, if the infrastructure is setup right, outages are very rare. Some of these were serving up massive loads at the time. I've done these build-outs a few times in my career and it isn't that hard to do, reliable software is more likely to be…

> It is as if the software industry has collectively forgotten how to run basic data center operations. Something that used to be a blue collar skill is now treated like arcane magic.

It's not arcane magic. It's undifferentiated toil that requires hiring for a different skill set than tech companies generally want to hire for. Of course when you get to a certain size it may make sense to take on this cost.

I want us to stop pretending individual actors lack the agency to make their own decisions, and they're all blind to how AWS is charging them a fortune for such simple things they can do themselves. You get value from AWS or you stop using AWS.

Re: Tell HN: AWS appears to be down again

#632

Earlier quoted context omitted.

all I've noticed is slack was a bit unreliable for a little bit, but i just carried on and otherwise ignored it. my world did not stop working.

My apartment block has a dialing system, that, instead if using a cale that goes to your apartment, relies on IP telephony and calls your mobile phone. It stos working if there is no internet, or your phone is out of battery, or you are not home but your wife is.

[deleted]

Re: Tell HN: AWS appears to be down again

#633
post #629

Earlier quoted context omitted.

> You would need to hire a hardware/ops specialist who can buy, spin up, and maintain physical machines. No. You rack the server, connect the cable to the switch press the on-switch, let it boot and then operate as you would with any other computer. Need a new server? Get the DCOps remote hands to Rack the server, cable it to the switch and press the on-switch. If you can install Windows 10, you can install Linux. >…

Configuring a server for high traffic, networking, various safety measures against fire/power, backups, connectivity to your external DBs, etc. etc. As a 'feature dev' of almost a decade, I could bungle my way through this and probably leave a bunch of security gaps open. I'm highlighting the fact that it is a specialized skill and a specific part of the stack that feature devs like me and many other early startup en…

> Configuring a server for high traffic, networking, various safety measures against fire/power, backups, connectivity to your external DBs, etc. etc.

All which you have to do on a cloud provider. Fire/Power are normally handled by the DC. You have to have someone with knowledge in the first place to operate that in the cloud and taking an application such as HAProxy is on the same skill level. Especially with vast fields of blogs you can find on the topic.

The cloud brings the instant "power-up" methodology and I would compromise, you could be right. If you want a perfective optimum cloud platform you would need a dedicated "cloud" engineer if you want to ensure security, connectivity etc. etc. But if your going to do that then again you might as well move in-house and hire a system admin, you've still got to operate your companies infrastructure. It's a moot point.

> The value to our customers for using cloud providers is the time saved on building infrastructure is instead spent on delivering value to the customer in the form of features/bug fixes.

I suppose this is a mixed area and what customers value varies. For me I value a service that uses it's own hardware rather than cloud. On the basis that they are willing to put the skill in to operate their own infrastructure rather than.

> I think you're vastly oversimplifying the task.

I don't think so. People seem to think that setting up what you can in the cloud is impossible on bare metal when really it isn't. What did devOp's do before the cloud providers? Amazon,Azure,X have only lassoed FOSS software, constructed a webGUI-admin panel and throw it as a service.

This is not to say Cloud doesn't has a purpose, otherwise it wouldn't exist today.

> it's something you're pretty familiar with so it makes sense that it's easy to you

While true, I won't disagree, I've been working in the SysAdmin field since 2009. But disagree as I started with very little knowledge and gained it through setting up such infrastructures. I've educated a few and those with very little knowledge of servers could setup an infrastructure that companies run in AWS.

Re: Tell HN: AWS appears to be down again

#634
post #196

Earlier quoted context omitted.

He asked it to demonstrate the point that uptime is trivial for one server with no traffic, and much harder at scale with auto scaling.

Then don't host with so many people? I don't think people care that AWS has other customers, they want their workload to work, if it doesn't: then that's a today issue.

I'm not sure your reply makes sense, seems like a non-sequiter. Individual companies might require thousands of servers and not want then running all the time. This means they either maintain thousands of servers on premises, or they use aws and auto scale. This has nothing to do with aws having other customers, and everything to do with only wanting to pay for and maintain as little as possible while being able to serve your applications.

Re: Tell HN: AWS appears to be down again

#635
post #294

Earlier quoted context omitted.

> The fewer API calls you need to make in-band with whatever throughput is generated via your customer demand, the better. Agreed; however, this is somewhat difficult to do correctly. There are all sorts of systems that might have hidden dependencies on managed services. e.g. AWS IAM roles will almost always be checked at some point if your services need to interact with AWS managed services. I think cloud providers…

It's not an AWS incentive thing really, it's a developer/consumer incentive thing. It's like the duality of modular code. If you want to manage one change in a lot of places, it's easiest to change it in the one module that everything else sources. But that means that one change to that module can take down everything. The alternative where you copy+paste the same change everywhere is the most resilient to failure, b…

I like the way you framed this, its a tradeoff mainly. You can build something fault-tolerant and highly-available, even during AWS outage events, but you have to give up a ton of the product offerings in their suite.

I've managed to stick to EC2/ELB and S3 as passive systems for the vast majority of what we build at my org (~90% of our stack). And for the most part, AWS failures are hitless for us as a result.

Re: Tell HN: AWS appears to be down again

#636
post #444

Earlier quoted context omitted.

There are setups where the UPS is designed to last long enough for generator spin up as well. I believe it's the most common setup if you have both. I assume spinning up the generators for very short-lived line power blips might be undesirable.

I was in a Bell Labs facility that had notoriously bad power. We occupied the building before the second main feed had been fully approved by the state and run to the building. Our main computer lab had a serial UPS that was online 100% of the time, though the inverters where under a very light load. If the mains even acted 'weird' (dips, bad power factor, spikes) the UPS jumped full on, and didn't revert to main pow…

>fully approved by the stat

Why in the world would the state need to be involved in this level of decision?

Re: Tell HN: AWS appears to be down again

#637
post #326

Earlier quoted context omitted.

You still need to pay someone to manage AWS infrastructure. It’s possible to save money using AWS, but things often get more expensive.

Of SMBs I’ve worked with, about 5% had a dedicated AWS engineer

Maybe non-tech businesses which I’m not very familiar with. But most startups absulutely have a dedicated DevOps engineer.

Re: Tell HN: AWS appears to be down again

#638

Earlier quoted context omitted.

Agree, and multi AZ is usually easy. IME with AWS and GCP the control plane is the same, the scaling works across AZ, bandwidth is free and latency is near zero. The level of effort to do that is simply ticking the right boxes at setup time IME.

Cross-AZ bandwidth is far from free and the biggest reason companies avoid it (IMO). Also latency is not near zero but I don't think that's the primary reason.

That's true. The moment your data leaves the region you start paying for egress and that can get expensive quickly. Still beats the crap out of multicloud, though :)

Re: Tell HN: AWS appears to be down again

#639

Earlier quoted context omitted.

Yes, I've seen issues that affected the entire region. In my specific case, I happened to have an ElastiCache cluster in the affected AZ that became unreachable (my fault for single AZ). But even now, I'm unable to create any new ElastiCache clusters in different AZs (which I wanted to use for manual failover). And there were a lot of errors on the AWS console during the outage. "almost unusable" is maybe exaggeratin…

Probably because you aren’t the only one trying to do that. The folks who successfully fail over a zone are the ones who have already automated the process and are running active/active configurations so everything is set up and ready to go.

No post body was provided.

Re: Tell HN: AWS appears to be down again

#640

Earlier quoted context omitted.

Amazon is claiming the failure is limited to a single AZ. Are you seeing failures for instances outside of that AZ? If not, how has this rendered "the entire region almost unusable"?

We've had alerts for packet loss and had issues in recovering region-spanning services (both AWS and 3rd party). Yes, some of these we should be better at handling ourselves, but... it's all very well to say "expect to lose an AZ" but during this outage it's not been physically possible to remove the broken AZ instances from multi-AZ services because we cannot physically get them to respond to or acknowledge commands…

No post body was provided.
Post reply on HN