Live data from Hacker News

LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active

longcat.chat

91–98 of 98 posts

Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active

#91

Earlier quoted context omitted.

> Which would ignore that jet fuel is going to be a multiple of the gasoline price. That doesn’t sound right. If my Duck Fu is any good, jet fuel is currently going due US$3.00 per gallon, avgas (leaded petrol) at $3.30, and gasoline at $2.88 gallon. There’s nothing much special about jet fuel, it’s just kerosene, same as RP1 (Rocket Propellant), heater fuel, and lamp oil you can buy from the hardware store, with a t…

in 2012 i owned a car that could run 100 octane fuel, and that was $9 a gallon. a few more octane and you get Jet-A minus the additives. according to jetfueltracker jetA is about $2 more than 87 octane right now and about $1 more than 93 octane. and still somehow cheaper than diesel. I'm not used to seeing jet fuel this cheap, luckily there's none near me to waste money on.

> in 2012 i owned a car that could run 100 octane fuel, and that was $9 a gallon. a few more octane and you get Jet-A minus the additives.

So? Diesel and petrol (gasoline) are different fuels. Comparing them by RON is irrelevant.

Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active

#92
post #71

Earlier quoted context omitted.

Yeah, but the default is extreme yap mode. I prefer "in a sentence" or "rows, not paragraphs." "Brutal" sounds like bait from an influencer trying to sell me a $400 course about how 4AM workouts will make me a millionaire.

I don't even know how we got here. This isn't that deeply represented in the training data. Is this what RLHF hath wrought? A new dialect of English based on corporatespeak and influencers, two heavy-hitting bullshitters?

Remember how, for SEO purposes, every food blog has to bs a multi-page story that everyone scrolls past to get to the 10 line recipe?

That's training data too; that's how we got here.

Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active

#94
post #90

Earlier quoted context omitted.

Specifically Dwarkesh couldn't understand that GPUs are not enough: it's GPUs plus multiple ecosystems to leverage them at massive scale during training vs inference. Instead of giving China open access to US controlled chips and creating a misalignment between labs that want to train a model on whatever is best, and hardware manufacturers that need labs to suffer the growing pains for their new ecosystems built from…

The Chinese ecosystem has not caught up; in fact, it's falling further behind, due to export restrictions on semiconductor manufacturing equipment. Even if America sold China all the chips Nvidia wants to, the CCP would still develop chips as quickly as possible as a matter of supply chain security.

While some years might pass until they will really catch up, that does not prevent them to find workarounds for their weaknesses.

For example, they recently have demonstrated a supercomputer faster than any of the US supercomputers.

Unlike the recent European supercomputers, which like the US supercomputers have been built by buying racks from HPE Cray, because China was not allowed to buy such things they have developed their own custom CPUs, designed in China, which have surpassed in throughput the AMD GPUs used in the fastest US supercomputers.

The Chinese CPUs match in memory bandwidth per socket the latest AMD MI355X GPUs (8 TB/s), while being significantly faster than the older AMD GPUs installed in the US supercomputers.

While the purpose of the new Chinese supercomputer is mainly for scientific/technical computing tasks that need high FP64 throughput, the CPUs used in it also have high enough BF16/INT8 performance and memory bandwidth and interconnection bandwidth (1.6 Tb/s directly from each CPU socket) to be able to train any big LLM.

So the evidence does not show China falling behind, but at least in certain directions they are already exceeding the performance of what they have been forbidden to buy.

For something like training a big LLM, the only disadvantage of the current Chinese devices is a lower energy efficiency, of only 65% to 70% of that of the best NVIDIA GPUs.

However that is not really a problem for China, as they have abundant cheap energy.

Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active

#95
post #16

Earlier quoted context omitted.

I am not a physicist but perhaps your question was leading more than you expected? I would take the question to pre-suppose I have an abundance of the stated material, ignoring practical realities of refinement. If I did have fully pure Pu-241, would that be a better fuel than U-235? Or stated another way, "If you could run a generator on gasoline or jet fuel, which one would you choose and why?" I would answer jet f…

If I did have fully pure Pu-241, would that be a better fuel than U-235? Also not a physicist, but I assume from the fact that the OP is asking the LLM this question to trip it up, the point is that U-235 is better even if you have an abundance of both. It's scarcity of Pu-241 leads to the lack of data in training, not that it's actually better.

I am lazy to search for a more authoritative source, but Wikipedia says that Pu-241 has a greater neutron absorption cross section than Pu-239 and the same probability of fission after absorbing a neutron. It would also generate a slightly higher amount of energy per fission event, than Pu-239.

This means that it would actually be a better fuel than Pu-239 for a fission reactor and only its scarcity prevents its use.

The only disadvantage versus Pu-239 is its short half-life, of 14 1/3 years. This means that it cannot be stored for a long time, so it must be consumed as a fuel soon after it is produced, to avoid losses.

Pu-239 has a lower delayed neutron fraction than U-235, which makes the control of a Pu-239 fission reactor more difficult.

But according to:

https://www-nds.iaea.org/sgnucdat/a6.htm

Pu-241 has almost the same delayed neutron fraction as U-235 (0.016 vs. 0.0162), so that is not a serious disadvantage for it.

Both Pu-239 and Pu-241 produce more neutrons per fission event than U-235. This simplifies some things, by allowing a reactor to work with less fuel or less-enriched fuel, but it complicates the control, because there is a greater risk of instabilities.

The truth is that an LLM cannot say which is a preferable fuel between U-235, Pu-239 and Pu-241. It would be possible to design fission reactors that work fine for any of these 3. The best choice depends on economical factors, not on technical feasibility factors.

The only real reason why Pu-241 will never be used is that its production yield when irradiating uranium with neutrons is too low in comparison with Pu-239, so it would be too expensive.

Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active

#96
post #90

Earlier quoted context omitted.

Specifically Dwarkesh couldn't understand that GPUs are not enough: it's GPUs plus multiple ecosystems to leverage them at massive scale during training vs inference. Instead of giving China open access to US controlled chips and creating a misalignment between labs that want to train a model on whatever is best, and hardware manufacturers that need labs to suffer the growing pains for their new ecosystems built from…

The Chinese ecosystem has not caught up; in fact, it's falling further behind, due to export restrictions on semiconductor manufacturing equipment. Even if America sold China all the chips Nvidia wants to, the CCP would still develop chips as quickly as possible as a matter of supply chain security.

Moving forward requires two parties with two very different financial incentives to cooperate in a way that harms the incentive of the other: the manufactures making the chips, and the companies using the chips to build AI.

Even if the Chinese government tells them to prioritize homegrown solutions, the A effort and players at labs are going to be on pushing the frontier, and the B effort/players goes to solving the teething problems

Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active

#97

I would love to see a 1.6T total with something like 3B active. I'm running an M4 Max and I'm still heavily bandwidth-limited -- I can hardly run anything at speed!

No, it won't fit in RAM and will read from SSD instead (10x+ slower).

Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active

#98
post #71

Earlier quoted context omitted.

I don't even know how we got here. This isn't that deeply represented in the training data. Is this what RLHF hath wrought? A new dialect of English based on corporatespeak and influencers, two heavy-hitting bullshitters?

Remember how, for SEO purposes, every food blog has to bs a multi-page story that everyone scrolls past to get to the 10 line recipe? That's training data too; that's how we got here.

It’s not only for that. The story provide them the copyright protection since recipes are not copyrighted
Post reply on HN