Live data from Hacker News

El Capitan: New supercomputer is the fastest

spectrum.ieee.org

71–80 of 106 posts

Re: El Capitan: New supercomputer is the fastest

#71
Fun facts, FFT was discovered back in 1965 based on the urgent necessity of discovering and detecting illegal nuke testing activities, just two years after the Partial Test Ban Treaty (PTBT) was signed in 1963 [1].

The first sentence statement in the article mentioning that United States and other nuclear powers committed to the Comprehensive Nuclear-Test-Ban Treaty in 1965 is wrong since the treaty was only signed in 1996 not in 1965 [2].

[1] The Algorithm That Almost Stopped The Development Of Nuclear Weapons:

https://www.iflscience.com/the-algorithm-that-almost-stopped...

[2] The Comprehensive Nuclear-Test-Ban Treaty:

https://www.ctbto.org/our-mission/the-treaty

Re: El Capitan: New supercomputer is the fastest

#72
post #42

Some may not want to hear this, but these “fastest supercomputer” list is now meaningless because all the Chinese labs have started obfuscating their progress. A while ago there were a few labs in China in top 10 and they all attracted sanctions / bad attention. Now no Chinese lab report any data now

They are in good company, with X, Meta, Microsoft and others not reporting theirs either. The basis for the ranking was a cumulative tracking of benchmark results that were required as part of commissioning bespoke computers. A contract would be written to buy a computer that could achieve a certain performance in operations per second, and in order to satisfy that the benchmarks were agreed to and codified in the co…

Microsoft has the #4 cluster on the top 500 list. Sure not everyone reports, still seems like a useful list to watch the trends in computing and in particular HPC.

Keep in mind the average hyperscalers cloud is not a particularly good setup for the top500. HPC tends towards more bandwidth, lower latency, and no virtualization.

Re: El Capitan: New supercomputer is the fastest

#73

Earlier quoted context omitted.

The DOE has entered the chat. (after the nuclear test ban treaty, they run a LOT of simulations)

Isn't that the open secret for El Cap? "Classified workloads" aka weapons sims.

Not a secret, from IEEE:

The NNSA—which oversees Lawrence Livermore as well as Los Alamos National Laboratory and Sandia National Laboratories—plans to use El Capitan to “model and predict nuclear weapon performance, aging effects, and safety,”

Re: El Capitan: New supercomputer is the fastest

#74

This is great but I absolutely love that poster of el capitan on the supercomputer racks ! Also TIL there is a list of top500 at https://www.top500.org/lists/top500/2024/11/

I've always loved these charts. The Numerical Wind Tunnel, #1 in 1993, achieved 124.2 gigaflops on the Linpack benchmark.

In comparison, the iPhone 15 Pro Max cellphone, released in 2023, delivers approximately 2150 gigaflops.

I once drew the chart backwards. I think my PC in 2013 would have been the fastest on Earth in 1990. And faster than every computer combined in about 1982.[0]

[0] might not be accurate

Re: El Capitan: New supercomputer is the fastest

#75
post #45

Earlier quoted context omitted.

A 64 bit float operation is >4X as expensive as a 16 bit float operation.

Agreed. However also note that if it was only matrix multiplies and no full transformer training, the performance of that Meta cluster would be closer to 16k PFlops/s, still much faster than the El Capitain performance measured on linpack and multiplied by 4. Other companies presumably cabled 100k H100s together, but they dont yet publish training data for their LLMs. It is good to have competition, I just didnt expe…

I'd expect linpack to be much closer to a user research application than training LLMs. My understanding of LLMs is that it's more about throughput and has a very predictable communication patterns, not latency sensitive, and bandwidth intensive.

Most parallel research, especially at this scale is more about different balance of operations to memory bandwidth, and much more worried about interconnect latency.

I wouldn't assume that just because various corporations have large training clusters that they could dominate HPC if they wanted to. Hyperscalers have dominated throughput for many years now, but HPC is a different beast.

Re: El Capitan: New supercomputer is the fastest

#76

Do super computers need proximity to other compute nodes in order to perform this kind of computations? I wonder what would happen if Apple offered people something like iCloud+ in exchange for using their idle M4 compute at night time for a distributed super computer.

The thing that sets these machines apart from something that you could set up in AWS (to some degree), or in a distributed sense like you're suggesting is the interconnect, how the compute nodes communicate. For a large system like El Capitan, you're paying a large chunk of the cost in connecting the nodes together, low latency, interesting topologies that ethernet, nor even Infiniband can get close to. Code that req…

I've seen nothing showing that slingshot has any particular advantage over IB for HPC. Sure HPE pushes slingshot (an HPE interconnect) over giving bags of money to Nvidia, but that's a business decisions. Eagle (the #4 cluster on the list) is Infiniband NDR.

I believe 306 of the top 500 clusters used Infiniband. Pretty sure the advance topologies like dragonfly are supported on IB as well as Slingshot. From what I can tell slingshot is much like ultra ethernet, trying to take the best of IB and ethernet and making a new standard. From what I can tell slingshot 11 latency is much like I got with omnipath/pathscale way back when dual core opterons were the cutting edge.

Re: El Capitan: New supercomputer is the fastest

#77

Fun facts, FFT was discovered back in 1965 based on the urgent necessity of discovering and detecting illegal nuke testing activities, just two years after the Partial Test Ban Treaty (PTBT) was signed in 1963 [1]. The first sentence statement in the article mentioning that United States and other nuclear powers committed to the Comprehensive Nuclear-Test-Ban Treaty in 1965 is wrong since the treaty was only signed i…

Another fun fact, the priority for detecting nuclear testing led to seisometers all over the planet. So the detection the exact position and nature of any disturbance on the planet became radically better. This was quite the boon to anyone interested in earthquakes, not only can the earth quake be detected in 2D, but accurately in 3D. The number and accuracy is enough you can see where on each fault is, the thickness of the crust, and the outline of subduction zones in 3d. Pretty crazy to see enough detail to see where plates enter the mantle and melts.

Said sensors can also track sonic booms from secret supersonic planes, but governments don't like to talk about that.

Re: El Capitan: New supercomputer is the fastest

#78
post #23

Noting here that 2700 quadrillion operations per second is less than the estimated sustained throughput of productive bfloat16 compute during the training of the large llama3 models, which IIRC was about 45% of 16,000 quadrillion operations per second, ie 16k H100 in parallel at about 0.45 MFU. The compute power of national labs has fallen far behind industry in recent years.

Any idea how that stacks up with GPT-4?

If I knew, I wouldn’t be able to disclose it :-)

Re: El Capitan: New supercomputer is the fastest

#80
post #27

Earlier quoted context omitted.

Are they always designing new nuclear bombs? Why the ongoing work to simulate?

Basically yes, we are always designing new nuclear bombs. This isn't done to increase yield, we've actually been moving towards lower yield nuclear bombs ever since the mid Cold War. In the 60s the US deployed the B41 bomb with a maximum yield of 25 megatons, making it the most powerful bomb ever deployed by the US. When the B41 was retired in the late 70s, the most powerful bomb in the US arsenal was the B53 with a…

> Another consideration is cost. Nuclear weapons are expensive to make, so a design that can get a high yield out of a small amount of fissile material is preferred. Maintenance, and the cost of maintenance, is also relevant. Will the weapon still work in 30 years, and how much money is required to ensure that?

I've seen speculation that Russia's (former Soviet) nuclear weapons are so old and poorly maintained that they probably wouldn't work. Not that anyone wants to find out.

Post reply on HN