Live data from Hacker News

Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

moonshotai.github.io

311–320 of 442 posts

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#311
post #236

Earlier quoted context omitted.

I suspect that the OpenRouter result originates from a quantized hosting provider. The difference compared to the direct API call from Moonshot is striking, almost like night and day. It creates a peculiar user and developer experience since OpenRouter enforces quantization restrictions only at the API level, rather than at the account settings level.

OpenRouter are proxying directly through to Moonshot - they're currently the only provider listed on https://openrouter.ai/moonshotai/kimi-k2-thinking/providers

That does include the Turbo endpoint, moonshotai/turbo. Add this to your prompt to only use the full-fat model:

-o provider '{ "only": ["moonshotai"] }'

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#312
post #139

Earlier quoted context omitted.

The latter. A reasoning model has been finetuned to use the scratchpad for intermediate results (which works better than just prompting a model to do the same).

I'd expect the same (fine tuning to be better than mere prompting) for most anything. So a model is or is not "a reasoning model" according to the extent of a fine tune. Are there specific benchmarks that compare models vs themselves with and without scratchpads? High with:without ratios being reasonier models? Curious also how much a generalist model's one-shot responses degrade with reasoning post-training.

> Are there specific benchmarks that compare models vs themselves with and without scratchpads? High with:without ratios being reasonier models?

Yes, simplest example: https://www.anthropic.com/engineering/claude-think-tool

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#313

Earlier quoted context omitted.

> What we’re going to see is as energy becomes a problem This is much more likely to be an issue in the US than in China. https://fortune.com/2025/08/14/data-centers-china-grid-us-in...

Disagree. Part of the reason China produces more power (and pollution) is due to China manufacturing for the US. https://www.brookings.edu/articles/how-do-china-and-america-... The source for China's energy is more fragile than that of the US. > Coal is by far China’s largest energy source, while the United States has a more balanced energy system, running on roughly one-third oil, one-third natural gas, and one-thir…

As counterpoints to illustrate Chinas current development:

* China has produced more PV panel capacity in the first half of this year than the US has installed, all in all, in all of its history

* China alone has installed PV capacity of over 1000 GW today

* China has installed battery electrical storage of about 100 GW / 300 GWh today and aims to have 180 GW in 2027

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#314

Earlier quoted context omitted.

And Europeans don't it because quite frankly, we're not really doing anything particularly impressive with AI sadly.

Europe should act and make its own, literal, Moonshot: https://ifiwaspolitical.substack.com/p/euroai-europes-path-t...

>Moonshot 1: GPT-4 Parity (2027) >Objective: 100B parameter model matching GPT-4 benchmarks, proving European technical viability

This feels like a joke... Parity with a 2024 model in 2027? The Chinese didn't wait, they just did it.

The timeline for #1 LLM is also so far into the future that it is entirely plausible that by 2031, nobody uses transformer based LLMs as we know them today anymore. For reference: The attention paper is only 8 years old. Some wild new architecture could come out in that time that makes catching up meaningless.

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#315

Earlier quoted context omitted.

> If OpenAI or Anthropic could squeeze the same output out of smaller GPUs and servers they'd be doing it for themselves. First, they do this; that's why they release models at different price points. It's also why GPT-5 tries auto-routing requests to the most cost-effective model. Second, be careful about considering the incentives of these companies. They all act as if they're in an existential race to deliver 'the…

> > If OpenAI or Anthropic could squeeze the same output out of smaller GPUs and servers they'd be doing it for themselves. > First, they do this; that's why they release models at different price points. No, those don't deliver the same output. The cheaper models are worse. > It's also why GPT-5 tries auto-routing requests to the most cost-effective model. These are likely the same size, just one uses reasoning and…

But they also squesed a 80% cut in O3 at some point, supposedly purely on inference or infra optimization

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#317
post #34

Earlier quoted context omitted.

I don't think this is the argument you want it to be, unless you're acknowledging the power of the Chinese government and their ability to suppress and destroy evidence. Even so there is photo evidence of dead civilians in the square. The best estimates we have are 200-10,000 deaths, using data from Beijing hospitals that survived. AskHistorians is legitimately a great resource, with sources provided and very strict…

The 10,000 number seems baseless The source for that is a diplomatic cable from the British ambassador within 48 hours of the massacre saying he heard it secondhand It would have been too soon for any accurate data which explains why it's so high compared to other estimates

Are you aware of any photographic evidence of civilian deaths inside Tiananmem Square?

I recently read a bit more about the Tiananmem Square incident, and I've been shocked at just how little evidence there actually is.

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#318

Earlier quoted context omitted.

Sure, but that's the point ... today's locally runnable models are a long way behind SOTA capability, so it'd be nice to see more research and experimentation in that direction. Maybe a zoo of highly specialized small models + agents for S/W development - one for planning, one for coding, etc?

> today's locally runnable models are a long way behind SOTA capability SOTA models are larger than what can be run locally, though. Obviously we'd all like to see smaller models perform better, but there's no reason to believe that there's a hidden secret to making small, locally-runnable models perform at the same level as Claude and OpenAI SOTA models. If there was, Anthropic and OpenAI would be doing it. There's…

I think SLM is developing very fast. A year ago, I couldn't have imagined a decent thinking model as Qwen, and now it seems full of promise

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#319
post #283
post #26

Earlier quoted context omitted.

Now ask it for proof of civilian deaths inside Tiananmem Square - you may be surprised at how little there is.

Huh? Please post the definitely proof you know to exist. Because it doesn't and that's one of the accusation toward the CCP, that they covered it up. It's funny that when the Israel government posted some photos of the Oct 7 massacres, people are very quick to point out that some seem staged. But some bloody photos that look like Tiananmem Square from the 80s is considered definite proof.

Israel has nothing to do with this. The horrific, indiscriminate genocide of Palestine and the creeping invasion of Lebanon and Syria are all happening right now in 4K. People nowadays know that you can't destroy thousands of vehicles with AK47's, and we've seen countless videos of Israeli military personnel admitting they killed many of their own people in a 'mass hannibal' event.

You do raise one good point however - propaganda in the time of Tiananmem was much, much easier before the advent of smartphones and the Internet. And also that Israel is really, really bad at propaganda.

Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model

#320
post #288

Earlier quoted context omitted.

Hard to be sure because the source of that information isn't known, but generally when people talk about training costs like this they include more than just the electricity but exclude staffing costs. Other reported training costs tend to include rental of the cloud hardware (or equivalent if the hardware is owned by the company), e.g. NVIDIA H100s are sometimes priced out in cost-per-hour.

Citation needed on "generally when people talk about training costs like this they include more than just the electricity but exclude staffing costs". It would be simply wrong to exclude the staffing costs. When each engineer costs well over 1 million USD in total costs year over year, you sure as hell account for them.

Table 1:

https://arxiv.org/html/2412.19437v2

Post reply on HN