Live data from Hacker News

Cheyenne Super Computer Auction

gsaauctions.gov

111–120 of 159 posts

Re: Cheyenne Super Computer Auction

#112
post #87
post #84

Earlier quoted context omitted.

Oh I'm speaking from experience with the SGI supercomputer blades. They're pretty wacky, 4x independent, dual cpu boards per blade and all sorts of weird connectors and cooling and management interfaces. Custom, centralized liquid cooling that requires a separate dedicated cooling rack unit and heat exchanger, funky power delivery with 3 phase, odd networking topologies, highly integrated cluster management software…

Yeah the licensing is often the stumbling block, unless you can just run some bog-standard linux on it. It sounds like this might be custom enough that it would be difficult (but I daresay we'll see a post in 5 years from someone getting part of it running after finding it on the side of the road).

Ultimately SGI was running Linux and AFAIK the actual hardware isn't using any secret sauce driver code, so yeah if you can get it powered on without it bursting in flames and get past the management locks you can probably get it working. It's definitely not impossible if you can somehow assemble the pieces.

Re: Cheyenne Super Computer Auction

#114
It's just not economical to run these given how power inefficient they are in comparison to modern processors.

This uses 1.7MW, or $6k per day of electricity. It would take only about four months of powering this thing to pay for 2000 5950X processors. Those would have a similar compute power to the 8000 Xeons in Cheyenne but they'd cost 1/4 the power consumption.

Re: Cheyenne Super Computer Auction

#115

Earlier quoted context omitted.

The Cheyenne numbers are 5.34 petaflops of *FP64*. The 5.6PF you quote for 18 A100's would be in BF16. Not comparable. The A100 can only do 9.746 TFLOPS in FP64. So you would need 548 A100's to match the FP64 performance of the Cheyenne.

AMD MI300x is 163.4 TFLOPS in FP64. 33 of them, which would also have 6,336TB of memory. I'll have way more than that in my next purchase order. It is really fun to build a super computer.

I used to build small clusters and use supercomputers and I can't imagine it's fun to build a super computer. It requires a massive infrastructure and significant employee base, and individual component failures can take down entire jobs. Finding enough jobs to keep the system loaded 24/7 while also keeping the interconnect (which was 15-20% of the total system cost) busy, and finding the folks who can write such jobs, is not easy. Even then, other systems will be constantly nipping at your heels with newer/cheaper/smaller/faster/cooler hardware.

Re: Cheyenne Super Computer Auction

#116
post #90

It's really hard to find a home for large, old, high-maintenance technology. What do you do with a locomotive, or a Linotype? They need a support facility and staff to be more than scrap. So they're really cheap when available. The Pacific Locomotive Association is an organization with that problem. About 20 locomotives, stored at Brightside near Sunol. They've been able to get about half of them working. It's all vo…

At the ill fated Portland TechShop I took woodworking classes from a retired gentleman, who professionally was a pattern maker for molding cast metal parts. This made his approach to woodworking really interesting. He had a huge array of freestanding sander machines, including a disc sander with more than a yard diameter. For anyone unfamiliar, pattern makers would make wooden model versions of parts that were to be…

Does 1/64" precision really mean anything in wood, where small fluctuations in air moisture can cause > 1/64" distortion? I guess it's OK if you stay within a climate controlled area.

Re: Cheyenne Super Computer Auction

#117

It's just not economical to run these given how power inefficient they are in comparison to modern processors. This uses 1.7MW, or $6k per day of electricity. It would take only about four months of powering this thing to pay for 2000 5950X processors. Those would have a similar compute power to the 8000 Xeons in Cheyenne but they'd cost 1/4 the power consumption.

If you can get 1.7MW service, then you are paying utility rates, or around 100-150 per MWh, or as you quoted 4-6k per day. In Seattle, running this off peak would cost one about 104$/hr before the other fees.

I would be neat if a subset of this could be made operational and booted once in awhile a computer history museum. I agree that it doesn't make sense to actually run it.

https://seattle.gov/city-light/business-solutions/business-b...

Re: Cheyenne Super Computer Auction

#118
post #4

>Components of the Cheyenne Supercomputer Installed Configuration: SGI ICE™ XA. E-Cells: 14 units weighing 1500 lbs. each. E-Racks: 28 units, all water-cooled Nodes: 4,032 dual socket units configured as quad-node blades Processors: 8,064 units of E5-2697v4 (18-core, 2.3 GHz base frequency, Turbo up to 3.6GHz, 145W TDP) Total Cores: 145,152 Memory: DDR4-2400 ECC single-rank, 64 GB per node, with 3 High Memory E-Cells…

Is it not "Personal" protective equipment? https://en.wikipedia.org/wiki/Personal_protective_equipment

You need PPE to protect your profession of moving stuff.

Re: Cheyenne Super Computer Auction

#119

The listing says that 1% of the nodes have RAM with memory errors. I assume this means hard errors since soft errors would just be corrected. Is this typical? Does RAM deteriorate over time?

Reading the whole paragraph, it sounds to me like they were accepting 1% errors rather than fixing the leaking cooling system.

Re: Cheyenne Super Computer Auction

#120
post #115

Earlier quoted context omitted.

AMD MI300x is 163.4 TFLOPS in FP64. 33 of them, which would also have 6,336TB of memory. I'll have way more than that in my next purchase order. It is really fun to build a super computer.

I used to build small clusters and use supercomputers and I can't imagine it's fun to build a super computer. It requires a massive infrastructure and significant employee base, and individual component failures can take down entire jobs. Finding enough jobs to keep the system loaded 24/7 while also keeping the interconnect (which was 15-20% of the total system cost) busy, and finding the folks who can write such job…

Thanks for the feedback. You make a lot of good points. I've built a 150,000 GPU system previously, but it was lower end hardware. It was a lot of fun to make it run smoothly with its own challenges.

It doesn't take a lot of employee's, we did the above on essentially two technical people. Those same two are working on this business.

Finding workloads/jobs is definitely going to be an interesting adventure, that said, the need for compute isn't going away. By offering hard to get hardware at reasonable rates and contract lengths, I believe we are in a good position on that front, but time will tell.

We are only buying the best of the best that we can get today. The plan is to continuously cycle out older hardware as well as not pick sides on one over another. This should help us keep pace with other systems.

Post reply on HN