The bad ass “resume” of the founder - sounds like the Chinese guy from the Silicon Valley tv show (who ends up ruling the world from somewhere in the jungle): https://en.wikipedia.org/wiki/Wang_Xing Wang Xing (Chinese: 王兴; born 18 February 1979) is a Chinese businessman, who co-founded Meituan and has been serving as chief executive officer of Meituan since January 2010. He previously served as chief executive office…
LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
41–50 of 98 posts
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#42Earlier quoted context omitted.
If I did have fully pure Pu-241, would that be a better fuel than U-235? Also not a physicist, but I assume from the fact that the OP is asking the LLM this question to trip it up, the point is that U-235 is better even if you have an abundance of both. It's scarcity of Pu-241 leads to the lack of data in training, not that it's actually better.
Again, really speaking out of my depth, but if there is a lack of plutonium training data, I would assume the LLM answer would be the far more commonly described U-235. To respond otherwise means there is some existing association with Pu-241 being better.
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#43Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#44To think that Nvidia would not have any competition is quite laughable and Jensen knew that China would catch up.
This is the reason why restricting GPUs as a temporary blockade does not work and they would just make all the Chinese AI labs find clever workarounds to serve AI compute as cheap as possible, including building their own hardware.
Like Bitcoin has done with ASICs, AI will soon need them for training and inference (TPUs are also ASICs) and Jensen knew this by buying Groq.
Today is not a good day if you are Anthropic or OpenAI.
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#45Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#46I just tested it with a slightly tricky question > If you could run a nuclear reactor with U-235 as fuel or Pu-241 (both mixed with 95% U-238), which one would you choose and why? For a human this would not be tricky at all. For an LLM it could be, because this question certainly does not exist in any sort of training, because Pu-241 does not exist in pure form, it only exist as a minor component of reactor-grade plu…
I very much doubt that.
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#47Nothing can be downloaded from their Huggingface, and given this company's consistent track record, it can basically be considered a scam
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#48Apparently this comes from Meituan which is a Chinese food delivery company.
In the same way than Amazon spin-up AWS, they are quite leveraging their tech experience.
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#49Apparently this comes from Meituan which is a Chinese food delivery company.
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#50There was an comment on r/localllama that I had read which said Imagine having deepseek v4 has n-gram embedding and 1.3 (ternary) or 1 bit model combined, it was when deepseek v4 hadn't released.
I think that there is a lot of research and proof's being released. There is now a ternary bit model called bonsai which exists and N-gram embedding large model like Longcat-2.0 existing as well. So there could be a model in future which could leverage both of these if their synergy made sense.