Kimi K3, and what we can still learn from the pelican benchmark
simonwillison.net
Kimi K3, and what we can still learn from the pelican benchmark
1–10 of 245 posts
Re: Kimi K3, and what we can still learn from the pelican benchmark
#2Re: Kimi K3, and what we can still learn from the pelican benchmark
#3I can't help but wonder where is the trend going? What will we have in five years? Maybe it will all have puttered out, and we will have moved to the next thing? Or maybe the prompt then will be "make a pelican ride a bicycle", and out will come the genetic code for a giant pelican with extremities suitable for a handle bar and pedals, and an inborn affinity to ride bicycles?
Re: Kimi K3, and what we can still learn from the pelican benchmark
#4Another day, another model and another pelican :-) I can't help but wonder where is the trend going? What will we have in five years? Maybe it will all have puttered out, and we will have moved to the next thing? Or maybe the prompt then will be "make a pelican ride a bicycle", and out will come the genetic code for a giant pelican with extremities suitable for a handle bar and pedals, and an inborn affinity to ride…
Re: Kimi K3, and what we can still learn from the pelican benchmark
#5Re: Kimi K3, and what we can still learn from the pelican benchmark
#6This is quite possibly reasoning-effort prompt which is injected before the opening token whenever you set a custom reasoning effort, see e.g. DeepSeek-V4 max mode prompt: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main...
Re: Kimi K3, and what we can still learn from the pelican benchmark
#7I'm starting to not trust any "benchmarks" when it comes to frontier models at least. As an example Sol feels the most "gets stuff done" but has zero taste, or any capability to surprise.
And for frontier models I go one step ahead and try to recreate a complex animation video, with the ability for the model to review its own work. And at this Fable is still the top one. Ex: https://www.youtube.com/watch?v=uDAeAuYyl0E (recreation of Claude announcement video) and https://www.youtube.com/watch?v=cSsVNtGPOIg (recreation of a fireship video). Sol did something similar but you can instantly tell its AI slop from very small things, and it just has no narrative or thought put into the writing.
https://mesmer.tools/benchmarks/ai-video-generation , I usually put basic ones here.
Re: Kimi K3, and what we can still learn from the pelican benchmark
#8Engineers get unbelievably silly about evaluating costs of things.
"The tokens are so expensive!" Oh my sweet child, how much would even the least capable human effort cost? This is what the executives properly understand that the programmers don't.
Re: Kimi K3, and what we can still learn from the pelican benchmark
#9My personal benchmark for new models has been to compare video making skills with something like remotion. Usually reveals if they have any "taste" or outside the box thinking. I'm starting to not trust any "benchmarks" when it comes to frontier models at least. As an example Sol feels the most "gets stuff done" but has zero taste, or any capability to surprise. And for frontier models I go one step ahead and try to…
Re: Kimi K3, and what we can still learn from the pelican benchmark
#10Another day, another model and another pelican :-) I can't help but wonder where is the trend going? What will we have in five years? Maybe it will all have puttered out, and we will have moved to the next thing? Or maybe the prompt then will be "make a pelican ride a bicycle", and out will come the genetic code for a giant pelican with extremities suitable for a handle bar and pedals, and an inborn affinity to ride…
> What will we have in five years? Maybe it will all have puttered out, and we will have moved to the next thing?
We will just have more of the same.