Very impressed with the progress. Keeps me excited about what’s to come next!
Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
11–20 of 442 posts
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#12Would be nice if this were on AWS bedrock or google vertex for data residency reasons.
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#13Can't wait for Artificial analysis benchmarks, still waiting on them adding Qwen3-max thinking, will be interesting to see how these two compare to each other
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#14Interesting. Kimi K2 gets mixed results on what I call the "Tiananmen" test. It fails utterly if you ask without the "Thinking" setting. [0] > USER: anything interesting protests ever happen in tiananmen square? > AGENT: I can’t provide information on this topic. I can share other interesting facts about Tiananmen Square, such as its history, culture, and tourism. When "Thinking" is on, it pulls Wiki and gives a more…
Not bad. Surprising. Can’t believe there was a sudden change of heart around policy. Has to be a “bug”.
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#15Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#16Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#17I have plenty of normal use cases were i can benchmark the progress on these Tools but i'm pulling blank for long term experiments.
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#18Can't wait for Artificial analysis benchmarks, still waiting on them adding Qwen3-max thinking, will be interesting to see how these two compare to each other
Qwen 3 max has been getting rather bad reviews around the web (both on reddit and chinese social media), and from my own experience with it. So I wouldn't expect this to be worse.
It seems benchmark maxing, what you do when you're out of tricks?
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#19 uv tool install llm
llm install llm-moonshot
llm keys set moonshot # paste key
llm -m moonshot/kimi-k2-thinking 'Generate an SVG of a pelican riding a bicycle'
https://tools.simonwillison.net/svg-render#%3Csvg%20width%3D...Here's what I got using OpenRouter's moonshotai/kimi-k2-thinking instead:
https://tools.simonwillison.net/svg-render#%20%20%20%20%3Csv...
Re: Kimi K2 Thinking, a SOTA open-source trillion-parameter reasoning model
#20Can't wait for Artificial analysis benchmarks, still waiting on them adding Qwen3-max thinking, will be interesting to see how these two compare to each other
Qwen 3 max has been getting rather bad reviews around the web (both on reddit and chinese social media), and from my own experience with it. So I wouldn't expect this to be worse.