Live data from Hacker News

Kimi K3, and what we can still learn from the pelican benchmark

simonwillison.net

51–60 of 245 posts

Re: Kimi K3, and what we can still learn from the pelican benchmark

#53

I would be surprised if pelican svgs are not part of the training corpus rn

If that were the case then it'd do a way better job. Think experienced artist level.

how would great pelicans make their way into the training set?

what they do have are many different pelicans and people helpfully rating them in the comments.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#54
post #43

Earlier quoted context omitted.

and that is bad because ?

the nature of the test was to see if the models can effectively compose an image of a novel concept outside the training set. If they are trained on it, it ceases to be an interesting test to some extent.

I would urge you to re-read the blog post you are commenting on. It pretty clearly explains how it is an interesting test independently of "see[ing] if the models can effectively compose an image of a novel concept outside the training set".

Re: Kimi K3, and what we can still learn from the pelican benchmark

#55
post #29

It's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website

Simon has stated a few times that he knows it’s possible that pelicans could be in the training sets. He also has other tests he doesn’t share publicly. He’s just a fan of pelicans.

From the article it doesn't even sound like he cares about pelicans at all, and doesn't think they are a good way to compare models anymore ... but people are used to seeing the test now, and it does serve as a common "hello world" unit of work.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#56

It's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website

Did you read the post? It's not even that long. He explicitly mentions this...

Respectfully, did you? The comment was specific to doubting the believe simonw has that labs are not training [0] specifically for this task, which is exactly what simonw wrote in the post [1], that it is a believe of his that they don't. He did not mention any kind of evidence or any piece of information that would indicate that the commenter didn't read the blog post.

Did you read either the post or the comment it was referencing?

On the note of training on SVGs, I have seen some labs models outperform when prompted for SVGs of certain animal and action combinations (pelican on bike, panda eating burger, etc.) compared to other similarly outlandish prompts for SVG output that are not part of widely reported benchmarks, even shared evidence one of the last times this came up on here.

[0] ... incredible Simon still believes ...

[1] I’m still not convinced that labs ....

Re: Kimi K3, and what we can still learn from the pelican benchmark

#57

It's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website

A person from Google famously put on her linkedin that her job was to optimize SVG for Gemini 3.0.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#58

It's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website

Imagine if we applied this train of logic to humans. "That artist saw a pelican at the beach once!" [cue the outrage] "He's not a real artist, he's a cheater and produces nothing original!"

We do. People who, for example, memorize question banks to pass certification tests without knowing the underlying material are equally frowned upon for not having the problem solving skills that they purport to. I'll leave the contrasts between LLMs and people to the well-written sibling comments.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#59
post #30

Do any of the vision models render the SVG and look at the result. Perhaps more importantly can they do that during reinforcement training. Learning how to critically analyse the appearance of what it generates would be quite useful. Manually feeding images back to models has been hilariously bad in the past which suggests that relating something it sees to something it wrote is not an ability it is very good at.

I imagine all vision models have to do this, this being html rendering, to be able to do well in web design.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#60
post #48

Earlier quoted context omitted.

Did you read the post? It's not even that long. He explicitly mentions this...

Clearly not. There's a subset of HN users who rush to post this same thing every single time.

Maybe it gets posted every time because besides a personal believe by the person popularising this "benchmark", there is no reason to assume that certain labs aren't intentionally training to game this and every other lab at least unintentionally gets improvements for this specific combination of animal and action because the internet is full of both good and bad examples, often ranked, which does inevitably become training data.

I have shared examples of certain models by certain labs doing far better on the pelican cycling vs other, similar prompts. Just operating on a feeling that labs don't optimise for this (as mentioned, even if they don't training data is filled with these) is not solid enough that criticism shouldn't be leveraged when it comes up.

Post reply on HN