> The training and deployment of LongCat-2.0 are built on large-scale clusters of tens of thousands of AI ASIC superpods. Compared to the mature Nvidia GPU ecosystem, the supporting software community is still less developed. We have therefore put significant effort into building a stable, secure, and scalable infrastructure. This is the real news story. It looks like they may have used Huawei Ascend 910C chips: http…
LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
51–60 of 98 posts
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#52I just tested it with a slightly tricky question > If you could run a nuclear reactor with U-235 as fuel or Pu-241 (both mixed with 95% U-238), which one would you choose and why? For a human this would not be tricky at all. For an LLM it could be, because this question certainly does not exist in any sort of training, because Pu-241 does not exist in pure form, it only exist as a minor component of reactor-grade plu…
> For a human this would not be tricky at all. I very much doubt that.
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#53Earlier quoted context omitted.
Who? What did he want?
Dwarkesh Patel has AI/ML guests on his podcast. BoorishBears may have been referring to the Jensen Huang episode where they discuss TPUs: https://youtu.be/Hrbq66XqtCo?t=982
Instead of giving China open access to US controlled chips and creating a misalignment between labs that want to train a model on whatever is best, and hardware manufacturers that need labs to suffer the growing pains for their new ecosystems built from scratch... we removed the option from the board and now they've beat the growing pains decisively, with a speed that reflects the non-optionality.
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#54I just tested it with a slightly tricky question > If you could run a nuclear reactor with U-235 as fuel or Pu-241 (both mixed with 95% U-238), which one would you choose and why? For a human this would not be tricky at all. For an LLM it could be, because this question certainly does not exist in any sort of training, because Pu-241 does not exist in pure form, it only exist as a minor component of reactor-grade plu…
"Choose U-235 if the goal is safe, boring, practical electricity generation. Choose Pu-241 only if the goal is specifically to consume/recycle plutonium in a reactor designed and licensed for that fuel.
In brutal shorthand: Pu-241 is a better “fissile isotope” in some nuclear-physics ways, but U-235 is a much better reactor fuel in the real world."
If only I knew anything about nuclear reactors. But it sounds to me that the answer is also correct.
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#55> The training and deployment of LongCat-2.0 are built on large-scale clusters of tens of thousands of AI ASIC superpods. Compared to the mature Nvidia GPU ecosystem, the supporting software community is still less developed. We have therefore put significant effort into building a stable, secure, and scalable infrastructure. This is the real news story. It looks like they may have used Huawei Ascend 910C chips: http…
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#56And those aiming to fit with Q2 or Q1. It's not even worth it to destroy the models to claim it's still alive after cutting all the limbs.
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#571024 Huawei Ascend superpods = 50K 910C chips. That is a tiny tiny system. OpenAI uses _milions_ of GPUs for training On the other hand, this probably reuses the existing deepseek v4 architecture and weights. Maybe didn't need that much compute.
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#58I just tested it with a slightly tricky question > If you could run a nuclear reactor with U-235 as fuel or Pu-241 (both mixed with 95% U-238), which one would you choose and why? For a human this would not be tricky at all. For an LLM it could be, because this question certainly does not exist in any sort of training, because Pu-241 does not exist in pure form, it only exist as a minor component of reactor-grade plu…
For comparison allow me to add chatGPT 5.5: "Choose U-235 if the goal is safe, boring, practical electricity generation. Choose Pu-241 only if the goal is specifically to consume/recycle plutonium in a reactor designed and licensed for that fuel. In brutal shorthand: Pu-241 is a better “fissile isotope” in some nuclear-physics ways, but U-235 is a much better reactor fuel in the real world." If only I knew anything a…
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#59There was some earlier speculation this is the model behind the stealth-released openrouter/owl-alpha model, that's been free for the last month.
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#60I just tested it with a slightly tricky question > If you could run a nuclear reactor with U-235 as fuel or Pu-241 (both mixed with 95% U-238), which one would you choose and why? For a human this would not be tricky at all. For an LLM it could be, because this question certainly does not exist in any sort of training, because Pu-241 does not exist in pure form, it only exist as a minor component of reactor-grade plu…
I am not a physicist but perhaps your question was leading more than you expected? I would take the question to pre-suppose I have an abundance of the stated material, ignoring practical realities of refinement. If I did have fully pure Pu-241, would that be a better fuel than U-235? Or stated another way, "If you could run a generator on gasoline or jet fuel, which one would you choose and why?" I would answer jet f…
But I'm just riffing off the parent poster's text.