Earlier quoted context omitted.
I don't think this is where you were going with your comment, but I'll mention this just because you're somewhat adjacent to a routine mistake in business: Uber is a people delivery company, but they've had a lot of bright engineers working for them on their infrastructure and software over the years, and that work has rippled out through the industry. Amazon (in VMWare's words) is "a company that sells books", and t…
And Google is the ad factory.
LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
61–70 of 98 posts
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#62> Both the full training run and the large-scale deployment are built entirely on AI ASIC superpods. Pretraining spans millions of accelerator-days across more than 35 trillion tokens, To think that Nvidia would not have any competition is quite laughable and Jensen knew that China would catch up. This is the reason why restricting GPUs as a temporary blockade does not work and they would just make all the Chinese AI…
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#63Earlier quoted context omitted.
For comparison allow me to add chatGPT 5.5: "Choose U-235 if the goal is safe, boring, practical electricity generation. Choose Pu-241 only if the goal is specifically to consume/recycle plutonium in a reactor designed and licensed for that fuel. In brutal shorthand: Pu-241 is a better “fissile isotope” in some nuclear-physics ways, but U-235 is a much better reactor fuel in the real world." If only I knew anything a…
Am i the only one feeling my soul being emptied every time i read another "brutal shorthand" or "honest take"?
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#64Earlier quoted context omitted.
Dwarkesh Patel has AI/ML guests on his podcast. BoorishBears may have been referring to the Jensen Huang episode where they discuss TPUs: https://youtu.be/Hrbq66XqtCo?t=982
Specifically Dwarkesh couldn't understand that GPUs are not enough: it's GPUs plus multiple ecosystems to leverage them at massive scale during training vs inference. Instead of giving China open access to US controlled chips and creating a misalignment between labs that want to train a model on whatever is best, and hardware manufacturers that need labs to suffer the growing pains for their new ecosystems built from…
The same scenario happens all the time when the US takes away something from China and China doubles down, gets into survival mode and then beats the US.
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#65Earlier quoted context omitted.
Yeah, for me it seems like a if you have to ask you can't run it" type question. In general the TL;DR is that anything above 35B needs hardware you buy basically only to run large LLMs, and if you have that hardware you don't need to ask the question.
That's simply not true. ~70B models can run fine (albeit somewhat slow) on consumer hardware with 64GB RAM. There are heavily quantized (Q1.x) models that are still usable on similar hardware. Granted recently there haven't been a lot of models of this size, but still, 35B isn't really the practical limit. 35B is mostly the limit if you're using consumer grade GPUs with limited RAM and need the model to run fast. Peo…
I'm saying that 64GB+ personal computers are vanishingly rare outside builds that were specifically done with AI in mind.
Gamers never saw the need for them, and even in software development 32GB was the standard until AI came along.
Yes, there were specialized use cases where they did exist, and yes, some people just wanted to max out the Macbooks but.. it was rare.
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#661024 Huawei Ascend superpods = 50K 910C chips. That is a tiny tiny system. OpenAI uses _milions_ of GPUs for training On the other hand, this probably reuses the existing deepseek v4 architecture and weights. Maybe didn't need that much compute.
I'm sure it also takes more compute effort to be at the frontier, rather than being able to distill and poach ideas from the frontier. No mistake that it's the same handful of labs taking turns at or near the frontier.
Anthropic claims deepseek has made 150K requests to their servers. Even if this number is correct, it takes far more requests to distill from a 3.2T model into a 1.6T model. 150K is closer to running a few benchmarks.
If anything, deepseek together with googles deepmind are the ones innovating while Anthropic and openAI are spending money and time on politics to try to hinder or ban competition.
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#67Earlier quoted context omitted.
Am i the only one feeling my soul being emptied every time i read another "brutal shorthand" or "honest take"?
It's genuinely depressing.
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#68There was some earlier speculation this is the model behind the stealth-released openrouter/owl-alpha model, that's been free for the last month.
Not speculation - they said it was.
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#69Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#70I asked a question with "Search" enabled, with the app set to English, and got results back in Chinese. Interesting view into how the LLM responds to its context.