Live data from Hacker News

Kimi K3, and what we can still learn from the pelican benchmark

simonwillison.net

81–90 of 245 posts

Re: Kimi K3, and what we can still learn from the pelican benchmark

#81
post #66

It's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website

The pelicans are still all rubbish. If they make it into the training set it doesn't help the models produce better pelicans, if anything it will make them perform worse!

Yes, I see your point.

Your pelican output is thus both in the training set and yet still outside the capability of the model architecture.

And so you are tracking both the capability of the training and also the capability of the querying!

When you receive your first outstanding pelican it will track a gain of capability.

(btw I first mentioned simonw-pelican-into-training-set in May 2025 on twitter.)

My 3D-egyptology-explainer showed a massive uplift for Kimi K3 and this tracks a much improved 3D capability.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#82
post #56

Earlier quoted context omitted.

Did you read the post? It's not even that long. He explicitly mentions this...

Respectfully, did you? The comment was specific to doubting the believe simonw has that labs are not training [0] specifically for this task, which is exactly what simonw wrote in the post [1], that it is a believe of his that they don't. He did not mention any kind of evidence or any piece of information that would indicate that the commenter didn't read the blog post. Did you read either the post or the comment it…

I'll note there's a difference between "pelicans on bikes aren't part of the training set" and "I’m still not convinced that labs are training for the benchmark".

I'm sure all sorts of crap pelican riding bicycle SVGs have ended up in the huge crawls of data that the labs feed into their pre-training steps.

What I'm questioning here is that there are labs who have sat down and deliberately tested and tweaked the performance for this particular task, independent of general model improvements.

The one exception here is Gemini, who have clearly invested a lot of effort in SVG tasks. I have no idea if my stupid benchmark influenced that decision!

Gemini have boasted about how good they are at pelicans riding bicycles, frogs on penny-farthings, giraffes driving a tiny car, ostriches on roller skates, turtles kickflipping skateboards, and dachshunds driving a stretch limousine. So if they trained for the test they did at least expand it a whole bunch! https://twitter.com/JeffDean/status/2024525132266688757

Re: Kimi K3, and what we can still learn from the pelican benchmark

#84

It's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website

A person from Google famously put on her linkedin that her job was to optimize SVG for Gemini 3.0.

SVG output is useful, though. I often ask whatever LLM I have open to generate placeholder icons whenever I need them.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#85
post #68
post #60

Earlier quoted context omitted.

Maybe it gets posted every time because besides a personal believe by the person popularising this "benchmark", there is no reason to assume that certain labs aren't intentionally training to game this and every other lab at least unintentionally gets improvements for this specific combination of animal and action because the internet is full of both good and bad examples, often ranked, which does inevitably become t…

> I have shared examples of certain models by certain labs doing far better on the pelican cycling vs other, similar prompts Please share those again! One of the things I'm most looking forward to is a lab producing a model that creates a really great pelican riding a bicycle and then a terrible sloth riding a skateboard (or whatever). I've not seen that myself yet.

Evidence in the other direction (that they're able to generalize) is that I can't think of any LLM currently that can't create usable (placeholder) SVG icons, I tried a bit before the pelican became popular and it was abysmal.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#86
post #66

It's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website

The pelicans are still all rubbish. If they make it into the training set it doesn't help the models produce better pelicans, if anything it will make them perform worse!

At this point I am simply interested in how much longer you're gonna ride this schtick

Re: Kimi K3, and what we can still learn from the pelican benchmark

#87
post #68
post #60

Earlier quoted context omitted.

Maybe it gets posted every time because besides a personal believe by the person popularising this "benchmark", there is no reason to assume that certain labs aren't intentionally training to game this and every other lab at least unintentionally gets improvements for this specific combination of animal and action because the internet is full of both good and bad examples, often ranked, which does inevitably become t…

> I have shared examples of certain models by certain labs doing far better on the pelican cycling vs other, similar prompts Please share those again! One of the things I'm most looking forward to is a lab producing a model that creates a really great pelican riding a bicycle and then a terrible sloth riding a skateboard (or whatever). I've not seen that myself yet.

Happy to, here one example where Grok 4 Fast, despite producing a fairly consistent pelican [0], did severely worse in a similarly outlandish scenario along with Haiku 4.5 and GPT-5 for context: https://news.ycombinator.com/item?id=45599403

> [...] a really great pelican riding a bicycle and then a terrible sloth riding a skateboard [...]

Happy to play ball. You made a blog post a few weeks back on one of the Qwen models with the eye-catching title "Qwen3.6-35B-A3B on my laptop drew me a better pelican than Claude Opus 4.7" [1].

Here is what Qwen3.6-35B-A3B via Openrouter provided for a sloth riding a skateboard: https://imgur.com/a/Dy8fvR5

Like Grok 4 Fasts attempt at a mushroom in a rowboat, it is barely recognisable as anything despite both Qwen3.6-35B-A3B and Grok 4 Fast having no issue with more popular (i.e. benchmarked) examples. Whether this is a case of training data being unsanitized or intentional benchmark targeted training, I cannot say, but it is the case.

And here is Opus 4.7, again via Openrouter: https://imgur.com/a/Qus1Enf

A massive delta in favour of Opus 4.7, despite the pelican Qwen3.6-35B-A3B produced being noticeably better as you rightly pointed out. What does that tell us? Whether intentional or not (with such deltas, I do have my suspicions), any eval with such a delta is clearly polluted and can not be a source of information, especially as its continued existence does hinge on you testing similar prompts in private as a sanity check, yet by your own admission never noticing the plainly apparent delta in quality. I specifically stuck with the skateboarding sloth too, to keep it as fair as possible and found this in less than 5 minutes...

I would not critique your use of this fun benchmark the way I tend to if I did not have evidence to back up my position, including private evals beyond SVGs that I can reliably use to point out major deviations between what a models claimed performance is according to major benchmarks vs the actual performance outside these known test cases.

I will also say that while I have a lot to be critical of regarding Anthropics modus operandi, especially how they present interesting findings like their j-space work, which I found was irresponsibly anthropomorphic in their reporting, especially as this wasn't a first in model interpretability, but mainly a leap due to being applied to a larger model, but of all the labs, they are the ones that never underperform my evals vs public ones and they appear to strictly keep their training data sanitised.

Happy to discuss public vs private evals and the merit of each if you'd like, I do appreciate your reporting in general but just think the SVG benches have become evidently polluted, which is also why even simple queries in my benchmarks are private. Just saw Thinking Machines Inkling model succeed in certain queries that neither Fable 5, nor GPT-5.6 Sol on any reasoning level managed, which I feel is valuable to truly gauge where we are at. Informs my work with models, my views of the industry and my assessment of the future these tools have, along with how to best implement them to enable better UX.

[0] https://simonwillison.net/2025/Sep/20/grok-4-fast/

[1] https://simonwillison.net/2026/Apr/16/qwen-beats-opus/

Re: Kimi K3, and what we can still learn from the pelican benchmark

#88
post #66

It's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website

The pelicans are still all rubbish. If they make it into the training set it doesn't help the models produce better pelicans, if anything it will make them perform worse!

I agree with that. I think, in particular, all the broken bike frames associated with "pelican on a bike" probably make it harder for LLMs to render correct bike frames.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#89
post #77

One thing i keep thinking: you only run the pelican once per model. Run the same model a few times and you get some different pelicans, so some of "this one is better" might just be which run you picked for it. Would love to see 8 runs per model side by side. I bet for two close models, the gap between runs is about as big as the gap between the models.

I've done versions in the past where I ran 3 and picked the best one. At some point I'd like to automate that with an LLM-as-a-judge (from the same model family) picking the "best" one to move forth in the competition. I built a whole ELO scoring mechanism a while back, described here: https://simonwillison.net/2025/Jun/6/six-months-in-llms/#ai-... I probably should spend some time on this now, even though the benchm…

[dead]

Re: Kimi K3, and what we can still learn from the pelican benchmark

#90
post #80

Earlier quoted context omitted.

Respectfully, the pelicans used to be an unrecognisable mess and now they’re unquestionably pelicans on bicycles, rendered poorly, from every model. In the same timescale, model capabilities across the board have only meaningfully improved in places where the labs are focusing their training efforts. Moreover, they have a uniform style, even though your prompt doesn’t ask for one. There's no model going rogue and pro…

You know what, that's actually something I hadn't considered before. There's definitely a bias towards a pelican cycling from left to right on a red bicycle against a blue sky and green grass. Blue sky and green grass aren't that surprising, but the color and direction are interesting. When I finally build the proper gallery I'll throw in a few other creature-vehicle combinations, and track some characteristics like…

In photography (and probably art in general), there's a composition "rule" to frame moving subjects from left to right.

So the direction may not be that interesting!

Post reply on HN