Live data from Hacker News

Kimi K3, and what we can still learn from the pelican benchmark

simonwillison.net

191–200 of 245 posts

Re: Kimi K3, and what we can still learn from the pelican benchmark

#191

The pelican benchmark is exactly what's wrong with hiring in technology. It's got nothing to do with what most people actually do when they're working - just like most job interviews which ask you to draw a pelican as their way of assessing you.

It's got nothing to do with what most people actually do when they're working.. AI companies claim their products are generalists though, and that they can do a good job on anything you give them, so you can't say what people will be doing with it. "Generate an SVG of an bird on a bicycle" is a corner case certainly but if a candidate interviewing for a role claims they can handle the corner cases then it's totally f…

Goodhart’s law is the problem, not the metric itself. Also LLMs do not have any visual generation skills, so its idea of a pelican looks like purely linguistic, unlike diffusion models. That we get decent results at all from an LLM outputting SVG files of random things is just nuts to me.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#192
post #90

Earlier quoted context omitted.

In photography (and probably art in general), there's a composition "rule" to frame moving subjects from left to right. So the direction may not be that interesting!

side scrolling video games were always moving from left to right

Well except for Jungle King

Re: Kimi K3, and what we can still learn from the pelican benchmark

#193
post #90

Earlier quoted context omitted.

In photography (and probably art in general), there's a composition "rule" to frame moving subjects from left to right. So the direction may not be that interesting!

The other thing to consider (as someone who frequently take a photos of their bike) the common direction has the drive side out! In cycling forums it is sacrilegious to post a photo of your bicycle without showing the drive side.

> the common direction has the drive side out

Took some searching and sleuthing to actually figure out what "drive side out" means, as I'm just a casual "from A to B" cyclist: apparently this is referring to the side the chainset, chainwheels and all those things are on.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#194
post #94

Earlier quoted context omitted.

I wonder if that changes in countries where the main language is written right to left?

It is. All over the Arab world, imagery in ads is “backwards” and I believe several companies will flip their ads horizontally, and UI localization involves flipping graphics.

Do the gears and stuff sit on the other side of the bike in the Arab world? Otherwise I'd expect cycling ads to still show a bike going from left to the right, considering https://news.ycombinator.com/item?id=48951828

Re: Kimi K3, and what we can still learn from the pelican benchmark

#196
Wild that we still haven't figured out how to make good benchmarks. What we really need is a way to properly quantify what makes a codebases architecture good, and then evaluate architecture of generated codebases, or evaluate refactors of existing ones.

Also, a way to evaluate a models ability to remove dead code, clean up slop, reorganize, etc.

None of the existing benchmarks test any of the things that truly matter. They were relevant when models struggled to one-shot functions, but we're so beyond that point right now, yet the industry has not kept up.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#198

The disconnection between pelican quality and overall model quality is interesting. I initially assumed that since pre-training is when a model gets its general skill that it happened around when RL started to really differentiate models. That is higher quality pre-trains result in higher quality pelicans, but RL is unlikely to touch pelican quality. However the fact that GLM 5.2 beats GPT 5.6 and Claude Fable puts a…

Correlation does not equal causation.

People seem to have forgotten this fact.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#199
post #66

Earlier quoted context omitted.

The pelicans are still all rubbish. If they make it into the training set it doesn't help the models produce better pelicans, if anything it will make them perform worse!

I agree with that. I think, in particular, all the broken bike frames associated with "pelican on a bike" probably make it harder for LLMs to render correct bike frames.

...even with the glut of pelicans, aren't there still far more images of actual bikes (with correct frames) available to train on?

Perhaps I'm underestimating the number of pelicans(?!)

Re: Kimi K3, and what we can still learn from the pelican benchmark

#200
post #141

Our answer to Pelican benchmark: https://playcode.io/blog/macbook-svg-benchmark

> Fable 5: Reasoned so long it exhausted the output budget before finishing the drawing. Lol

Wait, the user asked for a SVG of a pelican riding a bicycle. That doesn’t make sense, and I need to think about whether this is a legitimate request.

The user is asking to to generate an innocent and mundane graphic, possibly as part of a test.

But wait, pelicans cannot ride bicycles! A pelican is a water bird, and bicycles are designed to be ridden humans. Something alarming may be happening here, could this a jailbreaking attempt?

I need to reconsider and reread the user’s request, “make me a svg of a pelican riding a bicycle”. That is a perfectly innocent and legitimate task, as well as popular “benchmark” on social media communities, so I will continue. I need to continue to be on alert and watch out for potential jailbreaking attempts.

Post reply on HN