Pelican: https://tools.simonwillison.net/markdown-svg-renderer#url=ht... - rendered via the OpenRouter API: https://openrouter.ai/moonshotai/kimi-k3 95 input, 16,658 output = 25 cents! https://www.llm-prices.com/#it=95&ot=16658&ic=3&oc=15 (13,241 of those were reasoning tokens.) I think that's the most expensive pelican I've rendered through a Chinese model so far.
Hey Simon, I noticed one thing all LLMs are currently pretty bad at and maybe we could create a benchmark from it. Let an LLM play the role of a dungeon master and tell it to strictly stay in the script/story and only allow realistic player actions. You will notice that they are easily brought off track. E.g. - Tell the LLM that you as a player noticed a strange glow in an NPCs eyes -> the NPC becomes an enemy. - In…
Kimi K3: Open Frontier Intelligence
941–950 of 1001 posts
Re: Kimi K3: Open Frontier Intelligence
#942This might be the most impressive website generator demo I've seen: https://macos27.kimi.page Context from the person who prompted it: https://x.com/mweinbach/status/2077827886149439547
Re: Kimi K3: Open Frontier Intelligence
#943I switched to exclusively Chinese models, mostly Kimi, many months ago. I'll still ask Claude questions that require ambitious real-time web search / worldly knowledge, but for just about anything else, the Chinese models have been so good that I haven't looked back.
I think that those two statements are somewhat contradictory of one another. But in any case, I'm curious about your decision to still use Claude for some questions. Do you find the Chinese models have less worldly knowledge, and if so in what categories? Genuinely curious, I haven't had a chance to try them out very much.
The American chatbot apps are just very polished on live web search and tool use, and the models are very eager to do it. Chinese models are perfectly capable of that, but you need to bring your own MCPs (at least, if you want anything beyond WebFetch) and steer the model toward greater eagerness to use them.
Re: Kimi K3: Open Frontier Intelligence
#944Excited for the deepseek release this week (or at least they announced they'd release this week). Hopefully they also push even closer to SOTA.
That is exciting! I don't understand how DeepSeek can be so cheap with their cache pricing - ~0.003 usd / 1Mtok. 100x less than Kimi K3, or similar numbers against pretty much any other decently sized model to my knowledge. I've been using it whenever possible as even longer agent sessions cost few cents.
Re: Kimi K3: Open Frontier Intelligence
#945This might be the most impressive website generator demo I've seen: https://macos27.kimi.page Context from the person who prompted it: https://x.com/mweinbach/status/2077827886149439547
That is seriously impressive. Can I ask the prompt you gave it?
Re: Kimi K3: Open Frontier Intelligence
#946Earlier quoted context omitted.
An approach I like to help solving this is antagonistic or review agents. The first agent decides that eye glows turn NPCs into enemies, the second agent is fully dedicated to deciding if that is valid. If the review fails, it leaves notes and the original agent tries again.
If all or most LLMs are susceptible to failing this test, why would an LLM be a good means of evaluating performance on this test without being given a rubric?
Re: Kimi K3: Open Frontier Intelligence
#947This might be the most impressive website generator demo I've seen: https://macos27.kimi.page Context from the person who prompted it: https://x.com/mweinbach/status/2077827886149439547
Re: Kimi K3: Open Frontier Intelligence
#948US labs must be sweating bullets. Not on tech side, but on finance. They have a pile of debt and VC expectations that count on vast future profitability
They'll just lobby the government, feds will label them a security threat and ban them from corporate use. The American AI VC will breathe a huge sigh of relief and call it a day. Nothing says capitalism like writing laws to stop the competition needed to fuel innovation.
US is about 4% of world pop. And a lot of private use within those 4% will prioritize price.
Still a huge market with disproportionate spending power but from what I can tell OAI and friends need the outcome to be world conquering on the scale of google in search - global domination. Sub 4% ain't gonna cut it even if it's a thick 4%
Re: Kimi K3: Open Frontier Intelligence
#949The amount of scientific talent in China is astounding, also it's easier to be a fast follower than a innovator. Instead of limiting models and debating ethics like Anthropic, the edge lab should focus on lengthening the lead on China. The 2027 Chinese model could be one that beats the US.
Re: Kimi K3: Open Frontier Intelligence
#950This might be the most impressive website generator demo I've seen: https://macos27.kimi.page Context from the person who prompted it: https://x.com/mweinbach/status/2077827886149439547