Live data from Hacker News

AWS doesn't make sense for scientific computing

noahlebovic.com

161–170 of 281 posts

Re: AWS doesn't make sense for scientific computing

#161

A former colleague did his PHD in particle physics with a novel technique (matrix element method). I can't really explain it, but it is extremely CPU intensive. That working group did it on CERN's resources, and they had to borrow quotas from a bunch of other people. For fun they calculated how much it would have cost on AWS and came up with something ridiculous like 3 million euros.

The bigger experiments will routinely burn through tens of millions worth of computing. But 10 million euros isn't much for these experiments. The issue is that they are publicly funded: any country is much happier to build a local computing center and lend it to scientists than to fork the money over to an American cloud provider.

(The expensive part of these experiments is simulating billions of collisions and how the thousands of outgoing particles propagate though a detector the size of a small building. Simulating a single event takes around a minute on a modern CPU, and the experiments will simulate billions of events in a few months. If AWS is charging 5 cents a minute it works out to tens of millions easy.)

Re: AWS doesn't make sense for scientific computing

#162
post #33

Earlier quoted context omitted.

I think it really depends on the task. Where HIPAA violation is a real threat, the equation changes. And just for CYA purposes those projects can get pushed to a cloud. Which does not necessarily involve any attempts to make them any more secure, but this is a different topic. That said, many scientists are operating on premise hardware like this: some servers in a shared rack and an el-cheapo storage solutions with…

> And it works just fine for them. Until it doesn't because there's a fire or huge power surge or whatever. That's the point -- there's a lot of risk they're not taking into account, and by focusing on the "it works just fine for them", you're cherry picking the ones that didn't suffer disaster.

The point is, there's not need for everything to be 100% reliable in this context. If a fire destroys everything and their computational resources is unavailable for a few days, that's somewhat okay. Not ideal, but not a catastrophic loss either. Even data loss is no catastrophic - at worst it means redoing one or two weeks worth of computations.

Some sort of 80/20 principle is at works here. Most of the costs in professional cloud solutions comes from making the infrastructure 99.99% reliable instead of 99% reliable. It is totally worth it if you have millions of customers that expect a certain level of reliability, but a complete overkill if the worst case scenario from a system failure is some graduate student having to redo a few days worth of computations (which probably had to be redone several times anyway because of some bug in the code or something).

Re: AWS doesn't make sense for scientific computing

#163
post #146

What does the landscape look like now for "terraform for bare metal"?. Is ansible/chef still the main name in town? I just wanna netboot some lightweight image, set up some basic network discovery on a control plane, and turn every connected box into a flexible worker bee I can deploy whatever cluster control layer (k8s/nomad) on top of and start slinging containers.

I really like this description of how baremetal infrastructure should work, and this is where I think (shameless self promotion) Triton DataCenter[1] plays really well today on-prem.

PXE booted lightweight compute nodes with a robust API, including operator portal, user portal, and cli.

Keep an eye out for the work we are doing with Triton Linux + K8s. Very lightweight Triton Linux compute node + baremetal k8s deployments on Triton.

[1] https://www.tritondatacenter.com

Re: AWS doesn't make sense for scientific computing

#164
post #141

AWS is fantastic for scientific computing. With it you can: - Deploy a thousand servers with GPUs in 10 minutes, churn over a giant dataset, then turn them all off again. Nobody ever has to wait for access to the supercomputer. - Automatically back up everything into cold storage over time with a lifecycle policy. - Avoid the massive overhead of maintaining HPC clusters, labs, data centers, additional staff and train…

One thing to consider: I don't control my AWS account. I don't even have an AWS account in my professional life. I tell my IT department what I want. They tell the AWS people in central IT what they want. It's set up. At some point I get an email with login information. I email them again to turn it off. Do I hate this system? Yes. Is it the system I have to work with? Also yes. "AWS as implemented by any large insti…

[deleted]

Re: AWS doesn't make sense for scientific computing

#165

> Hardware is amortized over five years hardware running 100% won't last five years if hardware is not needed to be running 100% at full steam for five years, you can turn down instances on the cloud and you don't pay anything in 2 years you'll be stuck with the same hardware, while on the cloud you follow cpu evolution as it arrives to the provider all in all the comparison is too high level to be useful

I think you underestimate how long modern hardware can last. I have 8 to 12 year old PCs running non-stop, in a musty and damp basement.

Re: AWS doesn't make sense for scientific computing

#166
post #148

Earlier quoted context omitted.

I'm not talking about temporary outages, I'm talking about data loss. With AWS it's extremely easy to keep an up-to-date database backup in a different region. And it's great that you haven't personally encountered disaster, but of course once again that's cherry-picking. And it's not just a component overheating, it's the whole closet on fire, it's a broken ceiling sprinkler system going off, it's a hurricane, it's…

>"With AWS it's extremely easy to keep an up-to-date database backup in a different region" It is just as extremely easy on Hetzner or on premises

I would say even easier on prem as you don't need to wade 15 layers deep to do anything. Since I have moved to hosting my own stuff at my house, I have learned that connecting a monitor and keyboard to a 'sever' is awesome for productivity. I know where everything is, its fast as hell, and everything is locked down. Monitoring temps, adjusting and configuring hardware is just better in every imaginable way. Need more RAM, Storage, Compute? Slap those puppies in there and send it.

For home gamers like myself, it's has become a no brainer with advances in tunneling, docker, and cheap prices on Ebay.

Re: AWS doesn't make sense for scientific computing

#167

> Hardware is amortized over five years hardware running 100% won't last five years if hardware is not needed to be running 100% at full steam for five years, you can turn down instances on the cloud and you don't pay anything in 2 years you'll be stuck with the same hardware, while on the cloud you follow cpu evolution as it arrives to the provider all in all the comparison is too high level to be useful

> hardware running 100% won't last five years

Five year is a pretty typical amortisation schedule for HPC hardware. During my sysadmin days, of CPU, memory, cooling, power, storage, and networking, the only things that broke were hard disks and a few cooling fans. Disks were replaced by just grabbing a space and slotting it in, and fans were replaced by, well, swapping them out.

Modern CPUs and memory last a very long time. I think I remember seeing Ivy Bridge CPUs running in Hetzner servers in a video they put out, and they're still fine.

Re: AWS doesn't make sense for scientific computing

#168

> Hardware is amortized over five years hardware running 100% won't last five years if hardware is not needed to be running 100% at full steam for five years, you can turn down instances on the cloud and you don't pay anything in 2 years you'll be stuck with the same hardware, while on the cloud you follow cpu evolution as it arrives to the provider all in all the comparison is too high level to be useful

I think you underestimate how long modern hardware can last. I have 8 to 12 year old PCs running non-stop, in a musty and damp basement.

they don't just die, thermal paste dry up, fans gum up, gpu will live, but thermal throttling will mean it'll run at, say, 80%.

Re: AWS doesn't make sense for scientific computing

#169
post #138
post #95

Having worked for 2 of the largest cloud providers (1 of them beimg the largest) i have to say "The Cloud" just doesnt makes sense (maybe with the exception of cloud storage) yet for most use cases, this including start ups, small and, mid size companies its just way to expensive for the benefits it provides, it moves your hardware acquisitions /maintainance cost to development costs, you just think better/cheaper be…

Having worked in 3 startups that were AWS-first, I can say that you've learned the completely wrong lessons from your time at your cloud providers. Building on AWS has provided scale, security, and redundancy at a substantially lower cost than doing any on-prem solution (except for a shitty one strung together with lowendbox machines). The combined AWS bill for the three startups is less than the cost of an F5, even…

"The combined bill" during which time period?

1 Month, for sure. What about 1 year? Also did those companies required to provide any training or hiring to achieve that? Because you also need to add that to the cost comparison

If you are comparing one month bill agains 1 time purchase (which if is correctly chosen should not happen but once every 10 years at the earliest) for sure it will be cheaper. When it comes down to scalability, development and deployment, you should check your tech stack rather than your infrastructure. Kubernetees and containerization should easily take care of those with on premise hardware while also reducing complexity + you will no longer have to worry for off the chart network transit fees

Re: AWS doesn't make sense for scientific computing

#170

> Hardware is amortized over five years hardware running 100% won't last five years if hardware is not needed to be running 100% at full steam for five years, you can turn down instances on the cloud and you don't pay anything in 2 years you'll be stuck with the same hardware, while on the cloud you follow cpu evolution as it arrives to the provider all in all the comparison is too high level to be useful

> hardware running 100% won't last five years Five year is a pretty typical amortisation schedule for HPC hardware. During my sysadmin days, of CPU, memory, cooling, power, storage, and networking, the only things that broke were hard disks and a few cooling fans. Disks were replaced by just grabbing a space and slotting it in, and fans were replaced by, well, swapping them out. Modern CPUs and memory last a very lon…

if you expect downtime in the 5 year to replace fan and whatnot, you're not getting 100% of your money/perf back - and I didn't see that in the article.

if you have spares, spares need to be in the cost, and value lost to downtime stay minimal. but you have to include spares in the expenses. if you don't have spares, 1-2 day downtime is going to be a decent hit to value.

Post reply on HN