Live data from Hacker News

AWS doesn't make sense for scientific computing

noahlebovic.com

231–240 of 281 posts

Re: AWS doesn't make sense for scientific computing

#231

Earlier quoted context omitted.

Is a postdoc hacking a cluster something you have seen before? I am genuinely curious because I worked on a cluster owned by my university as an undergrad and everyone was kind of assumed to be trusted. If you had shell access on the main node you could run any job you wanted on the cluster. You could enhance security I just wonder about this threat model, that's an interesting one. I am sure it happens to be clear.

Yes. Probably not surprise that the postdoc was a PRC national. Very competent in their field of study, but also in this country with instructions from an APT group.

I had a feeling you would say that. I don't think this was part of our threat model until pretty recently. That was why I asked, because these stories aren't really internalized collectively yet I think and it's valuable to reconsider who we can trust and what threat actors might value.

Re: AWS doesn't make sense for scientific computing

#232
There is a lot of discussion about supercomputers in this article. I don't think public cloud providers can compete easily with traditional super computers because they are built for optimal processing of extremely large scale MPI workloads. Such workloads are not common so I expect that public cloud providers wouldn't bother optimizing for this niche use case (though I know they all have offerings). Also when you are only optimizing for a single variable (i.e. speed), you can make design choices that would be impossible to make in a more general situation.

Of course, not all scientific computing workloads require a traditional supercomputer. In fact, I suspect most do not.

Re: AWS doesn't make sense for scientific computing

#233
post #220

Having had the responsibility of providing HPC for a literal buildings full of scientists, I can say that it may be true that you can get computation cheaper with owned hardware, than in a cloud. Certainly pay as you go, individual project at a time processing will look that way to the scientist. But I can also say with confidence that the contest is far closer than they think. Scientists who make this argument almos…

Perspective from a computational biologist: Campus hosted HPC means the direct cost pressure is seen as IT staff and hardware related costs. Researchers are encouraged to use the available capacity. This is good. Externally-hosted HPC means every single compute job is seen as something that directly costs money. This negatively affects the quality of scientific output (research playfulness / creativity / focus on the…

>> seen as something that directly costs money.

Seen by whom? Dummies it seems. Maybe the dummies are the problem, not the computing or accounting models.

Re: AWS doesn't make sense for scientific computing

#234

Earlier quoted context omitted.

I'd much rather store HIPAA data on a server in my office or closet than worry I got all the IAM settings right. And if I fire someone, security makes sure they can't get in the building. You cannot say the same about the cloud. Yes, I know you can do cloud security right, but on prem security is just harder to mess up.

More than half of security penetrations in our institution (A medical center - research - med school complex) over the 8 years I worked there, ending in 2020, came through the research arm, even though Research accounted for no more than 10% of enabled servers in the infrastructure. And we're talking APT penetrations. They weren't looking for HIPAA data (although I used that example in my original post), they were lo…

I appreciate your perspective, but it seems the security team should be watching for reverse proxies, tunnels, and other firewall anomalies for on-prem hardware just as a normal course of biz. And if a PI installs a self-managed server, that really should not gum up the works.

All that being said, I have never worked at a place (or in a dept) whose threat profile made APT a real thing.

Re: AWS doesn't make sense for scientific computing

#235

Earlier quoted context omitted.

Author here! Spot instance pricing is better than on-demand, but it doesn't include data transfer, and it's still more expensive than on-prem/Hetzner/etc. Data transfer costs exceed the cost of the instance itself if you're transferring many TB off AWS. For of our more popular AWS instance types I use – a c5a.24xlarge, used for comparison in the post – the cheapest spot price over the past month in us-east-1 was $1.6…

FWIW, spot prices for c5a.24xlarge in us-east-2b and us-east-2c seem to have been under $0.92/hr for most of the last 3 months. So, assuming some flexibility on the choice of region, that would adjust your estimate to $0.92 / $1.69 * $1233.70/mo = $671.60/mo, which looks a lot more reasonable. Hopefully I did that math right. Data egress prices are definitely still ridiculous, I agree.

True! It sometimes drops even more, which definitely makes spot instances attractive. The r5.16xlarges had a ~80% discount and a Also, if all of scientific computing switched to AWS in order to exploit spot instance pricing, I don't think those market dynamics would stay the same.

As an aside, I've had trouble created large clusters of high memory instances in us-east-2. They might have increased capacity recently, though.

Re: AWS doesn't make sense for scientific computing

#236
A trend I've seen on HN over past few years is that people love showing off how they are able to save money by spending more of their own time, especially on infra/cloud things - if you calculate your own hourly rate correctly, it's oftentimes more costly to DIY than outsourcing to experts (e.g., managed cloud).

Re: AWS doesn't make sense for scientific computing

#237

Earlier quoted context omitted.

That only applies to government / government sponsored research. That is different from someone doing their own research

This is true, but to say it is "not a law", as you did, completely unqualified, is incorrect. If the research project is connected with a government grant (and many are) you need to pay attention to those laws. Many universities also have their own policies you need to follow, regardless. (Requiring informed consent and protecting people's privacy seems like a good thing.)

Let me repeat it another way. The law only restricts the actions of the government. Members of a university are not the government. Even if they took government money they could legally ignore all of that stuff. Worst case you will not get more funding from them in the future.

Re: AWS doesn't make sense for scientific computing

#238

Earlier quoted context omitted.

This is true, but to say it is "not a law", as you did, completely unqualified, is incorrect. If the research project is connected with a government grant (and many are) you need to pay attention to those laws. Many universities also have their own policies you need to follow, regardless. (Requiring informed consent and protecting people's privacy seems like a good thing.)

Let me repeat it another way. The law only restricts the actions of the government. Members of a university are not the government. Even if they took government money they could legally ignore all of that stuff. Worst case you will not get more funding from them in the future.

I believe you are technically correct, but that does not change the fact that universities have IRBs and will require reviews/approval if you are connected to that institution. You really think they're going to put their funding at risk? This seems very unlikely.

I found this article: https://journals.sagepub.com/doi/10.1177/1073110520917030

Re: AWS doesn't make sense for scientific computing

#239

Even as a big cloud detractor, I have to disagree with this. A lot of scientific computing doesn't need a persistent data center, since you are running a ton of simulations that only take a week or so, and scientific computing centers at big universities are a big expense that isn't always well-utilized. Also, when they are full, jobs can wait weeks to run. These computing centers have fairly high overhead, too, alth…

This is tangential to your point, but I’ll just mention that Azure has some properly specced out HPC gear: IB, FPGAs, the works. You used to be able to get time on a Cray XC with an Ares interconnect, but I never have occasion to use it, so I don’t know if you still can. They’ve been aggressively hiring top-notch HPC people for a while.

That's the Sentinel system. I worked on it when I was at Cray, and we did some covid stuff[1][2] with a researcher at UAH. We accelerated a docking code using some cool tech I created (in Perl, so there!) and some mods my teammates did to the queuing system.

The work won some award at SC20[3] (fka Supercomputing conference). I had considered submitting for the Gordon Bell prize, which had been specifically requesting covid work, though I thought the stuff we had done wasn't terribly sexy. We were getting ~250-500x better performance than single CPU runs.

Looking back over these, I gotta chuckle, as this (press releases) is pretty much the only time I'm called "Dr.". :D

Back to the OPs points, they are right. In most cases, cloud doesn't make sense for traditional HPC workloads. There are some special cases where it does, those tend to be large ephemeral analysis pipelines, as in bioinformatics and related fields. But for hardcore distributed (mostly MPI) code, running for a long time on a set of nodes interconnected with low latency networks, dedicated local nodes are the better economic deal.

During my stint at Cray, I was trying (quite hard) to get supercomputers, real classical ones, into cloud providers, or become a supercomputing cloud provider ourselves. The Met Office system is in Azure, is a Cray Shasta, but that was more of a special case. I couldn't get enough support for this.

Such is life. I've moved on. Still doing HPC, but more throughput maximized.

[1] https://www.uah.edu/science/departments/math/news/14954-uah-...

[2] A whole marketing writeup was done here https://www.hpe.com/us/en/newsroom/journey-to-accelerate-dru... . I tried very hard to correct the errors in the writeups. Sadly I wasn't successful.

[3] https://baudry-lab.uah.edu/news#h.121c63ayp0k0

Re: AWS doesn't make sense for scientific computing

#240
Having worked in the high performance computing field and in cloud hosted commercial applications, I can agree with the article but for entirely different reasons. The reason why some scientific computing shouldn't be done on AWS has to do with networking and latency between compute nodes. Supercomputers often use specialized networking hardware to get single digit microsecond latencies for data transfer between compute nodes and much higher network bandwidth than what you would normally find between EC2 nodes. This allows simulations to efficiently operate on really large data sets that span hundreds or thousands of nodes. The network topology between these nodes is often denser than a tree (think a 2D or 3D grid topology) and offers shorter paths between nodes.

All of this allows you to run code that you can't run in AWS unless it fits on one computer only. It's also way more expensive than clusters of commodity hardware.

For problems that are trivially parallelizable without much communication between nodes - I don't think that most universities can actually operate those cheaper than renting them from cloud computing services. A lot of these calculations don't take the staff to operate data centers, the cost of the building itself or the opportunity cost of using lots of space for this purpose vs something else into account. Economics of scale also kick in here. It's way cheaper per computer for AWS to admin a data center because they do this for orders of magnitude bigger data centers than your typical university.

Post reply on HN