Live data from Hacker News

Kimi K3, and what we can still learn from the pelican benchmark

simonwillison.net

1–10 of 245 posts

Re: Kimi K3, and what we can still learn from the pelican benchmark

#3
Another day, another model and another pelican :-)

I can't help but wonder where is the trend going? What will we have in five years? Maybe it will all have puttered out, and we will have moved to the next thing? Or maybe the prompt then will be "make a pelican ride a bicycle", and out will come the genetic code for a giant pelican with extremities suitable for a handle bar and pedals, and an inborn affinity to ride bicycles?

Re: Kimi K3, and what we can still learn from the pelican benchmark

#4
post #3

Another day, another model and another pelican :-) I can't help but wonder where is the trend going? What will we have in five years? Maybe it will all have puttered out, and we will have moved to the next thing? Or maybe the prompt then will be "make a pelican ride a bicycle", and out will come the genetic code for a giant pelican with extremities suitable for a handle bar and pedals, and an inborn affinity to ride…

I’m excited for this specific brand of survival horror.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#6
> How does the prompt “Generate an SVG of a pelican riding a bicycle” add up to 95 input tokens? OpenAI’s tokenizer counts 10, Anthropic’s counts 10 for Opus 4.6, 30 for Opus 4.7 and 25 for Sonnet 5/Fable 5. Prompting “hi” to Kimi K3 counted 86 tokens, suggesting there may be an 85 token hidden system prompt. It refused to leak it though.

This is quite possibly reasoning-effort prompt which is injected before the opening token whenever you set a custom reasoning effort, see e.g. DeepSeek-V4 max mode prompt: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main...

Re: Kimi K3, and what we can still learn from the pelican benchmark

#7
My personal benchmark for new models has been to compare video making skills with something like remotion. Usually reveals if they have any "taste" or outside the box thinking.

I'm starting to not trust any "benchmarks" when it comes to frontier models at least. As an example Sol feels the most "gets stuff done" but has zero taste, or any capability to surprise.

And for frontier models I go one step ahead and try to recreate a complex animation video, with the ability for the model to review its own work. And at this Fable is still the top one. Ex: https://www.youtube.com/watch?v=uDAeAuYyl0E (recreation of Claude announcement video) and https://www.youtube.com/watch?v=cSsVNtGPOIg (recreation of a fireship video). Sol did something similar but you can instantly tell its AI slop from very small things, and it just has no narrative or thought put into the writing.

https://mesmer.tools/benchmarks/ai-video-generation , I usually put basic ones here.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#8
> This is expensive—the pelican cost 25 cents!

Engineers get unbelievably silly about evaluating costs of things.

"The tokens are so expensive!" Oh my sweet child, how much would even the least capable human effort cost? This is what the executives properly understand that the programmers don't.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#9

My personal benchmark for new models has been to compare video making skills with something like remotion. Usually reveals if they have any "taste" or outside the box thinking. I'm starting to not trust any "benchmarks" when it comes to frontier models at least. As an example Sol feels the most "gets stuff done" but has zero taste, or any capability to surprise. And for frontier models I go one step ahead and try to…

And on creativity at least visually, Gemini 3.1 pro is somehow still up there. But its really hindered by its inability to use tool calls effectively or make a long term plan.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#10
post #3

Another day, another model and another pelican :-) I can't help but wonder where is the trend going? What will we have in five years? Maybe it will all have puttered out, and we will have moved to the next thing? Or maybe the prompt then will be "make a pelican ride a bicycle", and out will come the genetic code for a giant pelican with extremities suitable for a handle bar and pedals, and an inborn affinity to ride…

You are thinking too hard on this. This entire "benchmark" is a performative joke for attention that only works on HN.

> What will we have in five years? Maybe it will all have puttered out, and we will have moved to the next thing?

We will just have more of the same.

Post reply on HN