Earlier quoted context omitted.
Hey Simon, I noticed one thing all LLMs are currently pretty bad at and maybe we could create a benchmark from it. Let an LLM play the role of a dungeon master and tell it to strictly stay in the script/story and only allow realistic player actions. You will notice that they are easily brought off track. E.g. - Tell the LLM that you as a player noticed a strange glow in an NPCs eyes -> the NPC becomes an enemy. - In…
An approach I like to help solving this is antagonistic or review agents. The first agent decides that eye glows turn NPCs into enemies, the second agent is fully dedicated to deciding if that is valid. If the review fails, it leaves notes and the original agent tries again.
Kimi K3: Open Frontier Intelligence
871–880 of 1001 posts
Re: Kimi K3: Open Frontier Intelligence
#872Now, will they actually release the weights? Seems like Chinese model providers are slowly closing up, like Alibaba's Qwen 3.6 which did release weights (but not the biggest parameter count ones) and none for 3.7.
Re: Kimi K3: Open Frontier Intelligence
#873Pelican: https://tools.simonwillison.net/markdown-svg-renderer#url=ht... - rendered via the OpenRouter API: https://openrouter.ai/moonshotai/kimi-k3 95 input, 16,658 output = 25 cents! https://www.llm-prices.com/#it=95&ot=16658&ic=3&oc=15 (13,241 of those were reasoning tokens.) I think that's the most expensive pelican I've rendered through a Chinese model so far.
Hey Simon, I noticed one thing all LLMs are currently pretty bad at and maybe we could create a benchmark from it. Let an LLM play the role of a dungeon master and tell it to strictly stay in the script/story and only allow realistic player actions. You will notice that they are easily brought off track. E.g. - Tell the LLM that you as a player noticed a strange glow in an NPCs eyes -> the NPC becomes an enemy. - In…
But in general, I've experienced things similar to you. I've also found that LLMs are bad at subtext, e.g. hinting at an NPC being a werewolf or vampire.
Re: Kimi K3: Open Frontier Intelligence
#874This is the perfect opportunity for anyone who wants to boycott Anthopic and OpenAI!Anyone who cares about the future of Ai has no excuse to continue relying on companies that want to have a monopoly on intelligence.
Re: Kimi K3: Open Frontier Intelligence
#875Re: Kimi K3: Open Frontier Intelligence
#876So Chinese labs are driving essentially towards commodotized intelligence. Even if its a few months behind the US. Is this a classic 'commoditize my compliment' situation? They want to sell the hardware and infrastructure behind AI and make the software part not the value driver / moat? I can see it. But also even two Chinese labs sinking 100s of millions USD into training isn't exactly commoditization. It's still a…
Re: Kimi K3: Open Frontier Intelligence
#877Re: Kimi K3: Open Frontier Intelligence
#878Earlier quoted context omitted.
Google did this by creating Android to undercut Apple. Many US tech firms did this strategy in the 2000s and 2010s of supporting open source alternatives to their opponent's closed source money maker, to undercut the competition. Back then, it let open source have a big boost and we all benefited from that. Hopefully open weights models will do the same so that AI can be more democratized. I wouldn't want to live in…
Android wasn't created to undercut Apple. It was started long before the iPhone project was public. Android was created because Google felt their apps could be huge on mobile (correct) but that mobile operating systems of the time were painful and frustrating to develop for.
Re: Kimi K3: Open Frontier Intelligence
#879US labs must be sweating bullets. Not on tech side, but on finance. They have a pile of debt and VC expectations that count on vast future profitability
Re: Kimi K3: Open Frontier Intelligence
#880> Chip Design > As an early proof of concept, Kimi K3 designed a chip to serve a nano model built on its own architecture. In a single 48-hour autonomous run, K3 built, optimized, and verified the chip using open-source EDA tools on the Nangate 45nm library. Within 4 mm², the chip closes timing at 100 MHz and sustains over 8,700 tokens/s decode throughput in simulation, packing 1.46M standard cells, 0.277 MB of SRAM,…
Really feels like end game type stuff - AI designing its own next versions, designing its own chips, etc.. The advancement is slow, but fast - like a plant growing. We really are the boiling frogs now aren’t we? And the people with eyes wide open are us, and anyone that frequents this site really. Is this Milliways?