Live data from Hacker News

Kimi K3, and what we can still learn from the pelican benchmark

simonwillison.net

151–160 of 245 posts

Re: Kimi K3, and what we can still learn from the pelican benchmark

#152
post #82

Earlier quoted context omitted.

I'll note there's a difference between "pelicans on bikes aren't part of the training set" and "I’m still not convinced that labs are training for the benchmark". I'm sure all sorts of crap pelican riding bicycle SVGs have ended up in the huge crawls of data that the labs feed into their pre-training steps. What I'm questioning here is that there are labs who have sat down and deliberately tested and tweaked the perf…

Yes that's the obvious thing to do and why straightforward variants of known tests would also be treated as contaminated by anyone being even somewhat rigorous. I don't know why the standard is is to be sure that it is happening versus it being a plausible risk of making the results useless.

The pelican test has never pretended to be "rigorous". It's always openly been very much not that.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#153
post #80

Earlier quoted context omitted.

Respectfully, the pelicans used to be an unrecognisable mess and now they’re unquestionably pelicans on bicycles, rendered poorly, from every model. In the same timescale, model capabilities across the board have only meaningfully improved in places where the labs are focusing their training efforts. Moreover, they have a uniform style, even though your prompt doesn’t ask for one. There's no model going rogue and pro…

You know what, that's actually something I hadn't considered before. There's definitely a bias towards a pelican cycling from left to right on a red bicycle against a blue sky and green grass. Blue sky and green grass aren't that surprising, but the color and direction are interesting. When I finally build the proper gallery I'll throw in a few other creature-vehicle combinations, and track some characteristics like…

What's interesting is that given the fairly general and short in length prompt for the test, none of the models are attempting things like more discrete details of the bike. Such as showing V-brakes or dual 160mm disc rotors, rear derailleur, water bottle in a bottle cage, panniers, lights, saddlebag, the rider wearing a helmet, or other details that might be found on as vague a description as "a bicycle".

Re: Kimi K3, and what we can still learn from the pelican benchmark

#154

Earlier quoted context omitted.

> the pelicans used to be an unrecognisable mess and now they’re unquestionably pelicans on bicycles, rendered poorly, from every model You would not expect that to happen if the models trained on the unrecognizable mess, right? > model capabilities across the board have only meaningfully improved in places where the labs are focusing their training efforts And the labs clearly did focus on improving image rendering.…

I’m not suggesting Simon’s pelicans in the dataset are having a meaningful impact. I’m expecting that a company like ScaleAI has a product along the lines of “benchmax dataset: SimonW’s Pelican on Bikes test” which is a private curated series of well-drawn SVGs of animals riding vehicles for training and RL.

If they're benchmaxed on SVG pelicans then the outcome of that has still produced a surprisingly good generic SVG image generator.

Go invent your own random alternatives and the AI models have across the board gotten better over time. Insects playing sports, anthropomorphic fruits performing martial arts, wizards conjuring weapons of WWII, whatever you can imagine. I've tried a lot of these, well beyond what I think would be a reasonable thing to specifically train as combinations. If they have given it a corpus of SVG drawings it has learned to extrapolate.

(note: wizards conjuring a tank got me a surprise animated SVG with my Qwen 3.6 35B model)

Re: Kimi K3, and what we can still learn from the pelican benchmark

#155
post #80

Earlier quoted context omitted.

Respectfully, the pelicans used to be an unrecognisable mess and now they’re unquestionably pelicans on bicycles, rendered poorly, from every model. In the same timescale, model capabilities across the board have only meaningfully improved in places where the labs are focusing their training efforts. Moreover, they have a uniform style, even though your prompt doesn’t ask for one. There's no model going rogue and pro…

You know what, that's actually something I hadn't considered before. There's definitely a bias towards a pelican cycling from left to right on a red bicycle against a blue sky and green grass. Blue sky and green grass aren't that surprising, but the color and direction are interesting. When I finally build the proper gallery I'll throw in a few other creature-vehicle combinations, and track some characteristics like…

It'd be hard to fully compare, but I think a truly random "creature-vehicle" along side the pelican test would catch who's gaming and who's not.

I'd also enjoy the absurdism of "Herring on a pogostick"

Re: Kimi K3, and what we can still learn from the pelican benchmark

#156

One thing i keep thinking: you only run the pelican once per model. Run the same model a few times and you get some different pelicans, so some of "this one is better" might just be which run you picked for it. Would love to see 8 runs per model side by side. I bet for two close models, the gap between runs is about as big as the gap between the models.

[deleted]

Re: Kimi K3, and what we can still learn from the pelican benchmark

#157
post #28

Earlier quoted context omitted.

You mean the scale that AWS provides with Bedrock?

Bedrock needs to actually update their chinese models to the newest versions for this to matter.

And they need to support prompt caching, or customers stuck on Bedrock will still find the very expensive models from OpenAI and Anthropic prics-competitive with the Chinese ones.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#158
post #90
post #80

Earlier quoted context omitted.

You know what, that's actually something I hadn't considered before. There's definitely a bias towards a pelican cycling from left to right on a red bicycle against a blue sky and green grass. Blue sky and green grass aren't that surprising, but the color and direction are interesting. When I finally build the proper gallery I'll throw in a few other creature-vehicle combinations, and track some characteristics like…

In photography (and probably art in general), there's a composition "rule" to frame moving subjects from left to right. So the direction may not be that interesting!

side scrolling video games were always moving from left to right

Re: Kimi K3, and what we can still learn from the pelican benchmark

#159
post #90

Earlier quoted context omitted.

In photography (and probably art in general), there's a composition "rule" to frame moving subjects from left to right. So the direction may not be that interesting!

The other thing to consider (as someone who frequently take a photos of their bike) the common direction has the drive side out! In cycling forums it is sacrilegious to post a photo of your bicycle without showing the drive side.

Beat me to it - but I had the same thought. Most amateur and nearly all professional studio photographs of a bicycle will have it drive side out so I expect this plays some role in it.
Post reply on HN