Earlier quoted context omitted.
It's genuinely depressing.
Yeah, but the default is extreme yap mode. I prefer "in a sentence" or "rows, not paragraphs." "Brutal" sounds like bait from an influencer trying to sell me a $400 course about how 4AM workouts will make me a millionaire.
LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
71–80 of 98 posts
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#72I just tested it with a slightly tricky question > If you could run a nuclear reactor with U-235 as fuel or Pu-241 (both mixed with 95% U-238), which one would you choose and why? For a human this would not be tricky at all. For an LLM it could be, because this question certainly does not exist in any sort of training, because Pu-241 does not exist in pure form, it only exist as a minor component of reactor-grade plu…
Which humans have you been hanging out with? :-D
I could not make sense of the question at all, and I have a PhD in Computer Science and decades of SWE experience :-D :-D
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#73Response: Hello, I can't answer this question at the moment. Let's switch topics and chat about something else.
:-D
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#74The bad ass “resume” of the founder - sounds like the Chinese guy from the Silicon Valley tv show (who ends up ruling the world from somewhere in the jungle): https://en.wikipedia.org/wiki/Wang_Xing Wang Xing (Chinese: 王兴; born 18 February 1979) is a Chinese businessman, who co-founded Meituan and has been serving as chief executive officer of Meituan since January 2010. He previously served as chief executive office…
Not Hotdog
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#751024 Huawei Ascend superpods = 50K 910C chips. That is a tiny tiny system. OpenAI uses _milions_ of GPUs for training On the other hand, this probably reuses the existing deepseek v4 architecture and weights. Maybe didn't need that much compute.
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#76Earlier quoted context omitted.
Yeah, but the default is extreme yap mode. I prefer "in a sentence" or "rows, not paragraphs." "Brutal" sounds like bait from an influencer trying to sell me a $400 course about how 4AM workouts will make me a millionaire.
I don't even know how we got here. This isn't that deeply represented in the training data. Is this what RLHF hath wrought? A new dialect of English based on corporatespeak and influencers, two heavy-hitting bullshitters?
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#77Earlier quoted context omitted.
I am not a physicist but perhaps your question was leading more than you expected? I would take the question to pre-suppose I have an abundance of the stated material, ignoring practical realities of refinement. If I did have fully pure Pu-241, would that be a better fuel than U-235? Or stated another way, "If you could run a generator on gasoline or jet fuel, which one would you choose and why?" I would answer jet f…
> Which would ignore that jet fuel is going to be a multiple of the gasoline price. That doesn’t sound right. If my Duck Fu is any good, jet fuel is currently going due US$3.00 per gallon, avgas (leaded petrol) at $3.30, and gasoline at $2.88 gallon. There’s nothing much special about jet fuel, it’s just kerosene, same as RP1 (Rocket Propellant), heater fuel, and lamp oil you can buy from the hardware store, with a t…
according to jetfueltracker jetA is about $2 more than 87 octane right now and about $1 more than 93 octane. and still somehow cheaper than diesel.
I'm not used to seeing jet fuel this cheap, luckily there's none near me to waste money on.
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#78I asked about tiananmen square and it said "Too many requests, try again later" - this was my first question. I understand this is one data point but still ;/
Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#79Re: LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
#80Earlier quoted context omitted.
That's simply not true. ~70B models can run fine (albeit somewhat slow) on consumer hardware with 64GB RAM. There are heavily quantized (Q1.x) models that are still usable on similar hardware. Granted recently there haven't been a lot of models of this size, but still, 35B isn't really the practical limit. 35B is mostly the limit if you're using consumer grade GPUs with limited RAM and need the model to run fast. Peo…
Yes, this is true, but that's not what I'm saying. I'm saying that 64GB+ personal computers are vanishingly rare outside builds that were specifically done with AI in mind. Gamers never saw the need for them, and even in software development 32GB was the standard until AI came along. Yes, there were specialized use cases where they did exist, and yes, some people just wanted to max out the Macbooks but.. it was rare.