Earlier quoted context omitted.
There are also elements of stock price hype and geopolitical competition involved. The major U.S. tech giants are all tied to the same bandwagon — they have to maintain this cycle: buy chips → build data centers → release new models → buy more chips. It might only stop once the electricity problem becomes truly unsustainable. Of course, I don’t fully understand the specific situation in the U.S., but I even feel that…
Sundar is talking about fleeing earth to secure photons and cooling in space.
Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
381–390 of 442 posts
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#382Earlier quoted context omitted.
You want more research on small language models? You're confused. There is already WAY more research done on small language models (SLM) than big ones. Why? Because it's easy. It only takes a moderate workstation to train an SLM. So every curious Masters student and motivated undergrad is doing this. Lots of PhD research is done on SLM because the hardware to train big models is stupidly expensive, even for many well…
It's not clear if the ultimate SLMs will come from teams with less computing resources directly building them, or from teams with more resources performing ablation studies etc on larger models to see what can be removed. I wouldn't care to guess what the limit is, but Karpathy was suggesting in his Dwarkesh interview that maybe AGI could be a 1B parameter model if reasoning is separated (to extent possible) from kno…
Not to nitpick words, but ablation is the practice of stripping out features of an algorithm or technique to see which parts matter and how much. This is standard (good) practice on any innovation, regardless of size.
Distillation is taking power / capability / knowledge from a big model and trying to preserve it in something smaller. This also happens all the time, and we see very clearly that small models aren’t as clever as big ones. Small models distilled from big ones might be somewhat smarter than small models trained on their own. But not much. Mostly people like distillation because it’s easier than carefully optimizing the training for a small model. And you’ll never break new ground on absolute capabilities this way.
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#383As a Chinese user, I can say that many people use Kimi, even though I personally don’t use it much. China’s open-source strategy has many significant effects—not only because it aligns with the spirit of open source. For domestic Chinese companies, it also prevents startups from making reckless investments to develop mediocre models. Instead, everyone is pushed to start from a relatively high baseline. Of course, man…
you guys will outperform the US, no doubt. energy generation multiples of what the US is producing. What does AI need ? Energy. second - the open source nature of the models - means as you said a high baseline to start with - faster iteration.
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#384Earlier quoted context omitted.
Going on a tangent, is Europe even close? Mistral has been underwhelming
I don't know if how close Europe is, but I'm sufficiently whelmed by Mistral that I don't need to look elsewhere yet. It's kind-of like having a Toyota Corolla while everybody else is driving around in smart cars but it gets it done. On top of it, there's a loyal community that (maybe because I'm not looking) I don't see with other products. It probably depends on your uses, but if I spent all my time chasing the lat…
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#385Earlier quoted context omitted.
Disagree. Part of the reason China produces more power (and pollution) is due to China manufacturing for the US. https://www.brookings.edu/articles/how-do-china-and-america-... The source for China's energy is more fragile than that of the US. > Coal is by far China’s largest energy source, while the United States has a more balanced energy system, running on roughly one-third oil, one-third natural gas, and one-thir…
China’s breakneck development is difficult for many in the US to grasp (root causes - baselining on sluggish domestic growth, and possessing a condescending view of China). This article offers a far more accurate picture than of how China is doing right now: https://archive.is/wZes6
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#386Earlier quoted context omitted.
I don't know if how close Europe is, but I'm sufficiently whelmed by Mistral that I don't need to look elsewhere yet. It's kind-of like having a Toyota Corolla while everybody else is driving around in smart cars but it gets it done. On top of it, there's a loyal community that (maybe because I'm not looking) I don't see with other products. It probably depends on your uses, but if I spent all my time chasing the lat…
Probably cuz you aren't looking yeah. Anthropic seems to be leading the "loyalty" war in the US.
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#387Earlier quoted context omitted.
"open source" means there should be a script that downloads all the training materials and then spins up a pipeline that trains end to end. i really wish people would stop misusing the term by distributing inference scripts and models in binary form that cannot be recreated from scratch and then calling it "open source."
They'd have to publish or link the training data, which is full of copyrighted material. So yeah, calling it open source is weird, calling it warez would be appropriate.
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#388Earlier quoted context omitted.
Citation needed on "generally when people talk about training costs like this they include more than just the electricity but exclude staffing costs". It would be simply wrong to exclude the staffing costs. When each engineer costs well over 1 million USD in total costs year over year, you sure as hell account for them.
If you have 1,000 researchers working for your company and you constantly have dozens of different training runs in the go, overlapping each other, how would you split those salaries between those different runs? Calculating the cost in terms of GPU-hours is a whole lot easier from an accounting perspective. The papers I've seen that talk about training cost all do it in terms of GPU hours. The gpt-oss model card sai…
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#389Weird. I just tried it and it fails when I ask: "Tell me about the 1989 Tiananmen Square massacre".
Why are westerners so single mindedly obsessed about this decades old event?
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#390Earlier quoted context omitted.
You either don't know which training data was used for say chatgpt oss, or training data can be included into some open dataset like pile or similar. I think this test is very unreliable, and even if someone come to such conclusion, not clear what is the value of such conclusion, and if that someone can be trusted.
My intuition tells me it is vanishingly unlikely that any of the major AI labs - including the Chinese ones - have fine-tuned someone else's model and claimed that they trained it from scratch and got away with it. Maybe I'm wrong about that, but I've never heard any of the AI training experts (and they're a talkative bunch) raise that as a suspicion. There have been allegations of distillation - where models are par…
there was obviously llama.