Live data from Hacker News

Kimi K3: Open Frontier Intelligence

kimi.com

941–950 of 1001 posts

Re: Kimi K3: Open Frontier Intelligence

#941
post #111

Pelican: https://tools.simonwillison.net/markdown-svg-renderer#url=ht... - rendered via the OpenRouter API: https://openrouter.ai/moonshotai/kimi-k3 95 input, 16,658 output = 25 cents! https://www.llm-prices.com/#it=95&ot=16658&ic=3&oc=15 (13,241 of those were reasoning tokens.) I think that's the most expensive pelican I've rendered through a Chinese model so far.

Hey Simon, I noticed one thing all LLMs are currently pretty bad at and maybe we could create a benchmark from it. Let an LLM play the role of a dungeon master and tell it to strictly stay in the script/story and only allow realistic player actions. You will notice that they are easily brought off track. E.g. - Tell the LLM that you as a player noticed a strange glow in an NPCs eyes -> the NPC becomes an enemy. - In…

A knowledge graph of the D&D module should solve this.

Re: Kimi K3: Open Frontier Intelligence

#942

This might be the most impressive website generator demo I've seen: https://macos27.kimi.page Context from the person who prompted it: https://x.com/mweinbach/status/2077827886149439547

This is seriously impressive, it built a file system under the hood? I was able to create files, copy around and view/change it with terminal and UI. Absolutely insane

Re: Kimi K3: Open Frontier Intelligence

#943

I switched to exclusively Chinese models, mostly Kimi, many months ago. I'll still ask Claude questions that require ambitious real-time web search / worldly knowledge, but for just about anything else, the Chinese models have been so good that I haven't looked back.

I think that those two statements are somewhat contradictory of one another. But in any case, I'm curious about your decision to still use Claude for some questions. Do you find the Chinese models have less worldly knowledge, and if so in what categories? Genuinely curious, I haven't had a chance to try them out very much.

My use of "worldly knowledge" might have been sloppy. I really meant "real-time worldly knowledge".

The American chatbot apps are just very polished on live web search and tool use, and the models are very eager to do it. Chinese models are perfectly capable of that, but you need to bring your own MCPs (at least, if you want anything beyond WebFetch) and steer the model toward greater eagerness to use them.

Re: Kimi K3: Open Frontier Intelligence

#944

Excited for the deepseek release this week (or at least they announced they'd release this week). Hopefully they also push even closer to SOTA.

That is exciting! I don't understand how DeepSeek can be so cheap with their cache pricing - ~0.003 usd / 1Mtok. 100x less than Kimi K3, or similar numbers against pretty much any other decently sized model to my knowledge. I've been using it whenever possible as even longer agent sessions cost few cents.

it is ridiculous really. it is so cheap, that i can just run it basically 24/7.

Re: Kimi K3: Open Frontier Intelligence

#945
post #929

This might be the most impressive website generator demo I've seen: https://macos27.kimi.page Context from the person who prompted it: https://x.com/mweinbach/status/2077827886149439547

That is seriously impressive. Can I ask the prompt you gave it?

It's not me, click on the X link to see Max Weinbach's post. He's moved on to seeing if it will do a bare-metal OS in Swift.

Re: Kimi K3: Open Frontier Intelligence

#946

Earlier quoted context omitted.

An approach I like to help solving this is antagonistic or review agents. The first agent decides that eye glows turn NPCs into enemies, the second agent is fully dedicated to deciding if that is valid. If the review fails, it leaves notes and the original agent tries again.

If all or most LLMs are susceptible to failing this test, why would an LLM be a good means of evaluating performance on this test without being given a rubric?

Because evaluating a performance is different task from creating a performance. Someone who plays an instrument badly often has a different perception from someone who has to listen to it. ;)

Re: Kimi K3: Open Frontier Intelligence

#947

This might be the most impressive website generator demo I've seen: https://macos27.kimi.page Context from the person who prompted it: https://x.com/mweinbach/status/2077827886149439547

I'm pretty annoyed with how fast this feels. Wish MacOS was this fast launching things.

Re: Kimi K3: Open Frontier Intelligence

#948
post #875

US labs must be sweating bullets. Not on tech side, but on finance. They have a pile of debt and VC expectations that count on vast future profitability

They'll just lobby the government, feds will label them a security threat and ban them from corporate use. The American AI VC will breathe a huge sigh of relief and call it a day. Nothing says capitalism like writing laws to stop the competition needed to fuel innovation.

It's a sound prediction on regulation, but I don't think that's enough.

US is about 4% of world pop. And a lot of private use within those 4% will prioritize price.

Still a huge market with disproportionate spending power but from what I can tell OAI and friends need the outcome to be world conquering on the scale of google in search - global domination. Sub 4% ain't gonna cut it even if it's a thick 4%

Re: Kimi K3: Open Frontier Intelligence

#949
post #713

The amount of scientific talent in China is astounding, also it's easier to be a fast follower than a innovator. Instead of limiting models and debating ethics like Anthropic, the edge lab should focus on lengthening the lead on China. The 2027 Chinese model could be one that beats the US.

The idea that china doesn’t innovate is ridiculous
Post reply on HN