Live data from Hacker News

HPC Systems Special Offer: Two A64FX Nodes in a 2U for $40k

anandtech.com

41–46 of 46 posts

Re: HPC Systems Special Offer: Two A64FX Nodes in a 2U for $40k

#41

An A64FX CPU has 48 cores. 1 CPU/node means 96 cores on 2 nodes for $40k and ~6TFLOPS (3 per). Meanwhile the 128 thread Threadripper 3990X costs $3.6k, and benchs put it at 1.5TFLOPS in Linpack, which is less but not 10x less, and I don't know what the benchmark basis for the A64FX value is, so I suspect it's closer, especially since it says theoretical peak performance. ... Am I missing something?

Well the A64FX "deal" is for those interested in building HPC clusters out of them. As such people people are interested in useful work / total cost to own.

Your 3990X system has dramatically less memory bandwidth with 4 64 bit memory channels good for a peak bandwidth of around 100GB/sec, which is 10% of the A64FX.

Your $3,500 price likely doesn't include an infiniband card, or a port on an IB leaf switch, or the port on spine IB switch, and you are getting quite a bit less memory bandwidth and floating point than the A64FX system. As a result you end up buying more threadripper nodes, paying for more power, more rack space, more cooling, more IB spine switches, more IB cables, more IB leaf switches, and potentially even a larger building ... just to hit the same performance.

Generally to get close to the memory bandwidth or flops with an x86-64 system you end up using an accelerator, like say the Nvidia A100. Problem is that the nvidia GPUs can only run CUDA aware programs, which rules out a significant chunk of HPC workloads.

The attraction of the A64FX is that it's efficient, has impressive flops per watt (which means cheaper cooling, racks, buildings), impressive memory bandwidth, and will run any python, perl, fortran, C, C++, go, java, code you throw at it.

For large clusters this isn't just a nice to have, it's a huge driver of real world performance / price. So much so that the top 4 supercomputers in the world are not x86-64. Even #5 uses a Intel Xeon to run the OS and do I/O and uses an accelerator (Matrix-2000) to do the heavy lifting. #6 and #7 are similar, but using a Nvidia cards. #8 is the first pure x86-64 cluster and is pretty small in comparison, 18 times smaller than #1.

Even #8 (Frontera) being 18 times smaller, makes it MUCH easier to scale codes that run across the entire cluster. Despite that huge advantage, Frontera scaled lipack to the entire cluster with 60% efficiency. The #1 Fugaku using the A64FX scaled at 80% efficiency, that's a pretty large real world advantage.

Keep in mind that Linpack is a pretty limited benchmark, heavily optimized, and not particularly memory or network intensive. Many real world codes will more heavily exercise the system and would likely show a larger performance differential than linpack.

Imagine you have 20,000 watts a rack and 100 racks. Sure you could use cheap nodes, but your total performance would be less... and that's before you throw in problems like scaling to 100 racks at 60% efficiency instead of 80%. When asking for your budget it will be that much harder if you only perform well on CUDA codes. Not that there's not room for cheaper node clusters out there, but non-x86-64 systems do seem to have a significant advantage on larger clusters.

Also general HPC devel/test systems often come with substantial support to enable users to tune their codes to a new platform. I wouldn't assume that if you want 100 racks of them that you'd pay $20k each, in fact the hardware you'd likely use isn't even the same hardware. The A64FX nodes used in clusters like Fugaku have higher density nodes that depend on water cooling, have more cores, and have a much better interconnect.

Re: HPC Systems Special Offer: Two A64FX Nodes in a 2U for $40k

#42
post #33

Earlier quoted context omitted.

Yeah, these computers can be spec with 32gb of HBM also.

no they can't, but unless Fujitsu or you produce some benchmark, to show me that 150GB of HBM with 250 cores beats the aforementioned supermicro I doubt that this is a cost-effective solution for general HPC (we have a KNL system here at the cluster, it isn't either – or I am to dumb to compile the software I use there...). Like KNL it's probably king for some specialized tasks, which is totally fine and is probably…

The #1 cluster on the top 500 list would like to have a chat.

Actually it looks like it's just the opposite of what you describe. Fujitsu build a CPU that's very good at running a wide variety of applications. Unlike x86-64's that need accelerators for good FP or memory bandwidth per node, the A64FX does not need CUDA applications. Plain old fortran is just fine.

Keep in mind the $40k includes support for porting applications, it doesn't mean that if you want 1000 of them that you'll have to pay $20k each.

Re: HPC Systems Special Offer: Two A64FX Nodes in a 2U for $40k

#43
post #28

An A64FX CPU has 48 cores. 1 CPU/node means 96 cores on 2 nodes for $40k and ~6TFLOPS (3 per). Meanwhile the 128 thread Threadripper 3990X costs $3.6k, and benchs put it at 1.5TFLOPS in Linpack, which is less but not 10x less, and I don't know what the benchmark basis for the A64FX value is, so I suspect it's closer, especially since it says theoretical peak performance. ... Am I missing something?

Fugaku supercomputer had a total cost of 1B$, and 160k cpus, so the cost was less than 6500 per chip, or 3600$ per 1.5 TFlops

Yes, $6250 per node, but that cost includes storage, compute nodes, racks, IB switches, GigE switches, consoles, PDUs, etc. At least 1/3rd of that $1B went to storage/network/infrastructure and likely much more.

Also note that the #1 on the top 500 list is using the a64fx and scales with 80% efficiency. The #8 on the list is a pure intel (no accelerators) and is 14 times smaller, and only scales with 60% efficiency. That's a pretty impressive feat since good scaling gets harder as the cluster increases in size.

Re: HPC Systems Special Offer: Two A64FX Nodes in a 2U for $40k

#44

An A64FX CPU has 48 cores. 1 CPU/node means 96 cores on 2 nodes for $40k and ~6TFLOPS (3 per). Meanwhile the 128 thread Threadripper 3990X costs $3.6k, and benchs put it at 1.5TFLOPS in Linpack, which is less but not 10x less, and I don't know what the benchmark basis for the A64FX value is, so I suspect it's closer, especially since it says theoretical peak performance. ... Am I missing something?

A64FX has HBM2 RAM. This means it will be relatively low (32GBs RAM), but extremely high performance RAM. Literally the highest-bandwidth RAM in the market, directly wired onto the chip itself over an interposer (PCB is too slow! Direct interposer connections only mm away from the cores). Your implication is correct however. x86 is more in line with "typical" consumers, and even businesses, who need this kind of comp…

HBM2 is totally tame in terms of signalling rate. I think it's like 2 GT/s, while GDDR6X is at 20 GT/s (both per pin). PCBs aren't too slow for HBM2, a silicon interposer is just the only practical way to route a memory interface using something like 7000 pins directly between wafers. If HBM2 chips were using regular packaging and pitches suitable for PCB use, each stack would have to be a huge package with thousands of balls.

Re: HPC Systems Special Offer: Two A64FX Nodes in a 2U for $40k

#45
post #42
post #33

Earlier quoted context omitted.

no they can't, but unless Fujitsu or you produce some benchmark, to show me that 150GB of HBM with 250 cores beats the aforementioned supermicro I doubt that this is a cost-effective solution for general HPC (we have a KNL system here at the cluster, it isn't either – or I am to dumb to compile the software I use there...). Like KNL it's probably king for some specialized tasks, which is totally fine and is probably…

The #1 cluster on the top 500 list would like to have a chat. Actually it looks like it's just the opposite of what you describe. Fujitsu build a CPU that's very good at running a wide variety of applications. Unlike x86-64's that need accelerators for good FP or memory bandwidth per node, the A64FX does not need CUDA applications. Plain old fortran is just fine. Keep in mind the $40k includes support for porting app…

yes, because running on the whole number #1 cluster is just recompiling your non-trivial software. no, it's not (just take a look at QuantumEspresso, which is widely used, not too bad, but does not scale tooo god with non-Fortran and QE experts testing). And as I said, this is a fine price, I just doubt that it's worth anyones attention who is not building an upper 6-figures to some millions cluster. And unless you actually port your application with Fujitsu (which incurs also personell cost) to take advantage of the arch, you will just fare better if you buy a bog-standard EPYC or even Skylake-Refresh, which "just works the same way it did before"©

Re: HPC Systems Special Offer: Two A64FX Nodes in a 2U for $40k

#46
post #45
post #42

Earlier quoted context omitted.

The #1 cluster on the top 500 list would like to have a chat. Actually it looks like it's just the opposite of what you describe. Fujitsu build a CPU that's very good at running a wide variety of applications. Unlike x86-64's that need accelerators for good FP or memory bandwidth per node, the A64FX does not need CUDA applications. Plain old fortran is just fine. Keep in mind the $40k includes support for porting app…

yes, because running on the whole number #1 cluster is just recompiling your non-trivial software. no, it's not (just take a look at QuantumEspresso, which is widely used, not too bad, but does not scale tooo god with non-Fortran and QE experts testing). And as I said, this is a fine price, I just doubt that it's worth anyones attention who is not building an upper 6-figures to some millions cluster. And unless you a…

Not so sure, many of the impressive A64FX numbers were published with no source code changes. Even large complex codes like WRF (a weather simulator) "just works" and was faster than a dual socket Xeon 8168 (each Xeon 8168 is $5680 each list).

For another F90 benchmark (Himeno) they got 4 times faster than a dual Xeon 8168 AND 1.1x faster than a Tesla V100 running the CUDA version of the same benchmark. IMO that's pretty amazing, no source code changes and you get BETTER than GPU performance without having to rewrite your code.

So sure a standard Epyc or Skylake refresh will run your code and "just work", the a64fx opens up the possibility of significant performance upgrades with no source code changes. If your code isn't currently rewritten for CUDA getting 10x the bandwidth and 4x the CPU performance would be really attractive.

Post reply on HN