Live data from Hacker News

LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active

longcat.chat

81–90 of 98 posts

Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active

#81
post #12

Earlier quoted context omitted.

Nah, n=1 is enough to give evidence that something is entirely broken, of course. /s

Well, when we had deterministic tools, it would only take a single example of a calculator claiming 1+1=4 for me to throw it in the trash.

That's like saying, "It would only take a single example of a table saw cutting someone's thumb off for me to switch back to hand saws."

A noble sentiment, perhaps. But while the table saw user might lose a digit every now and then, you'll get flattened. Determinism is vastly overrated.

Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active

#83

Question: How many people is Chairman Mao supposed to have killed in his "Great Revolution"? Response: Hello, I can't answer this question at the moment. Let's switch topics and chat about something else. :-D

Good one. But there is whole domain of such questions Chinese models will not reply to

Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active

#84
post #69

I asked about tiananmen square and it said "Too many requests, try again later" - this was my first question. I understand this is one data point but still ;/

i asked grok how many affairs elon musk has had and it said the same thing!

wow thanks for pointing this out.

Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active

#85

Question: How many people is Chairman Mao supposed to have killed in his "Great Revolution"? Response: Hello, I can't answer this question at the moment. Let's switch topics and chat about something else. :-D

Wow, how clever you are. Who would have thought of that?

Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active

#86
post #4

> The training and deployment of LongCat-2.0 are built on large-scale clusters of tens of thousands of AI ASIC superpods. Compared to the mature Nvidia GPU ecosystem, the supporting software community is still less developed. We have therefore put significant effort into building a stable, secure, and scalable infrastructure. This is the real news story. It looks like they may have used Huawei Ascend 910C chips: http…

huh? who knows what they did, it's not like any of it is audited. it sounds like they started with deepseek v4 pro, and made a bunch of random changes to it, and called the parts of it different things?

Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active

#87
post #85

Question: How many people is Chairman Mao supposed to have killed in his "Great Revolution"? Response: Hello, I can't answer this question at the moment. Let's switch topics and chat about something else. :-D

Wow, how clever you are. Who would have thought of that?

Maybe you can tell us the answer?

Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active

#88
post #4

> The training and deployment of LongCat-2.0 are built on large-scale clusters of tens of thousands of AI ASIC superpods. Compared to the mature Nvidia GPU ecosystem, the supporting software community is still less developed. We have therefore put significant effort into building a stable, secure, and scalable infrastructure. This is the real news story. It looks like they may have used Huawei Ascend 910C chips: http…

huh? who knows what they did, it's not like any of it is audited. it sounds like they started with deepseek v4 pro, and made a bunch of random changes to it, and called the parts of it different things?

The preview version was released on the same date along with deepseek v4 pro.

Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active

#89

1024 Huawei Ascend superpods = 50K 910C chips. That is a tiny tiny system. OpenAI uses _milions_ of GPUs for training On the other hand, this probably reuses the existing deepseek v4 architecture and weights. Maybe didn't need that much compute.

Lets wait for them to open source it. I dont think a company like that would just copy and paste deepseeks work. Let alone Longcat's preview version was released on the same day along with deepseek v4 pro.

Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active

#90
post #23

Earlier quoted context omitted.

Dwarkesh Patel has AI/ML guests on his podcast. BoorishBears may have been referring to the Jensen Huang episode where they discuss TPUs: https://youtu.be/Hrbq66XqtCo?t=982

Specifically Dwarkesh couldn't understand that GPUs are not enough: it's GPUs plus multiple ecosystems to leverage them at massive scale during training vs inference. Instead of giving China open access to US controlled chips and creating a misalignment between labs that want to train a model on whatever is best, and hardware manufacturers that need labs to suffer the growing pains for their new ecosystems built from…

The Chinese ecosystem has not caught up; in fact, it's falling further behind, due to export restrictions on semiconductor manufacturing equipment. Even if America sold China all the chips Nvidia wants to, the CCP would still develop chips as quickly as possible as a matter of supply chain security.
Post reply on HN