Live data from Hacker News

LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active

longcat.chat

41–50 of 98 posts

Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active

#41
post #30

The bad ass “resume” of the founder - sounds like the Chinese guy from the Silicon Valley tv show (who ends up ruling the world from somewhere in the jungle): https://en.wikipedia.org/wiki/Wang_Xing Wang Xing (Chinese: 王兴; born 18 February 1979) is a Chinese businessman, who co-founded Meituan and has been serving as chief executive officer of Meituan since January 2010. He previously served as chief executive office…

Not Hotdog

Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active

#42
post #16

Earlier quoted context omitted.

If I did have fully pure Pu-241, would that be a better fuel than U-235? Also not a physicist, but I assume from the fact that the OP is asking the LLM this question to trip it up, the point is that U-235 is better even if you have an abundance of both. It's scarcity of Pu-241 leads to the lack of data in training, not that it's actually better.

Again, really speaking out of my depth, but if there is a lack of plutonium training data, I would assume the LLM answer would be the far more commonly described U-235. To respond otherwise means there is some existing association with Pu-241 being better.

That's 2.5% more neutrons, surely that must be better!

Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active

#44
> Both the full training run and the large-scale deployment are built entirely on AI ASIC superpods. Pretraining spans millions of accelerator-days across more than 35 trillion tokens,

To think that Nvidia would not have any competition is quite laughable and Jensen knew that China would catch up.

This is the reason why restricting GPUs as a temporary blockade does not work and they would just make all the Chinese AI labs find clever workarounds to serve AI compute as cheap as possible, including building their own hardware.

Like Bitcoin has done with ASICs, AI will soon need them for training and inference (TPUs are also ASICs) and Jensen knew this by buying Groq.

Today is not a good day if you are Anthropic or OpenAI.

Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active

#46

I just tested it with a slightly tricky question > If you could run a nuclear reactor with U-235 as fuel or Pu-241 (both mixed with 95% U-238), which one would you choose and why? For a human this would not be tricky at all. For an LLM it could be, because this question certainly does not exist in any sort of training, because Pu-241 does not exist in pure form, it only exist as a minor component of reactor-grade plu…

> For a human this would not be tricky at all.

I very much doubt that.

Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active

#47
post #32

Nothing can be downloaded from their Huggingface, and given this company's consistent track record, it can basically be considered a scam

Meituan published LongCat Flash last year: https://huggingface.co/meituan-longcat/LongCat-Flash-Chat So their track record seems non-scammy so far. Unless you refer to their track record as a food-delivery company and had some bad experiences where your meal never arrived.

Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active

#48

Apparently this comes from Meituan which is a Chinese food delivery company.

It's mostly a conglomerate nowadays (e.g the list of subsidiaries in Wikipedia is huge https://en.wikipedia.org/wiki/Meituan).

In the same way than Amazon spin-up AWS, they are quite leveraging their tech experience.

Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active

#50
The N-gram embedding model thing is absolutely crazy. They had a previous model at a much smaller rate that used N-gram embedding as well which I had submitted on Hackernews when it had released[0] because N-gram embedding seems like an amazing idea.

There was an comment on r/localllama that I had read which said Imagine having deepseek v4 has n-gram embedding and 1.3 (ternary) or 1 bit model combined, it was when deepseek v4 hadn't released.

I think that there is a lot of research and proof's being released. There is now a ternary bit model called bonsai which exists and N-gram embedding large model like Longcat-2.0 existing as well. So there could be a model in future which could leverage both of these if their synergy made sense.

[0]: https://news.ycombinator.com/item?id=46803687

Post reply on HN