>$40k gets you almost-Opus GLM 5.2 is "almost Opus," and it needs at least 8xH200s for comfortable inference (so it's closer to $400k than $40k). They suggest using this modified model: >A REAP-pruned (≈22% of experts removed), Int8-mix NVFP4 quantized version of GLM-5.2, ≈594B parameters. I wonder how it behaves in practice outside of benchmarks. Qwen3.6, even at 6-bit quantization, often gets stuck in loops while r…
"GLM 5.2 is "almost Opus," and it needs at least 8xH200s for comfortable inference ..." What is the behavior if one were to run GLM 5.2 with only a single H200 ? Would it fail to run at all, or would it just run so slowly as to be unusable ? I would like to prove out the build, and concept, of a SOTA model locally, but then backfill the rest of the GPUs in 18-24 months when they cost significantly less ...
going to need you to sit down for this one...