Live data from Hacker News

Kimi K3, and what we can still learn from the pelican benchmark

simonwillison.net

91–100 of 245 posts

Re: Kimi K3, and what we can still learn from the pelican benchmark

#91
post #66

Earlier quoted context omitted.

The pelicans are still all rubbish. If they make it into the training set it doesn't help the models produce better pelicans, if anything it will make them perform worse!

Respectfully, the pelicans used to be an unrecognisable mess and now they’re unquestionably pelicans on bicycles, rendered poorly, from every model. In the same timescale, model capabilities across the board have only meaningfully improved in places where the labs are focusing their training efforts. Moreover, they have a uniform style, even though your prompt doesn’t ask for one. There's no model going rogue and pro…

> the pelicans used to be an unrecognisable mess and now they’re unquestionably pelicans on bicycles, rendered poorly, from every model

You would not expect that to happen if the models trained on the unrecognizable mess, right?

> model capabilities across the board have only meaningfully improved in places where the labs are focusing their training efforts

And the labs clearly did focus on improving image rendering.

> they have a uniform style

SVG output from LLMs always looks like that. It looked that way from the beginning; no LLM ever produced a watercolor when asked for SVG output. They all render the prompted element centered in the picture. They all tend to draw things going from left to right, and so on.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#92
post #90
post #80

Earlier quoted context omitted.

You know what, that's actually something I hadn't considered before. There's definitely a bias towards a pelican cycling from left to right on a red bicycle against a blue sky and green grass. Blue sky and green grass aren't that surprising, but the color and direction are interesting. When I finally build the proper gallery I'll throw in a few other creature-vehicle combinations, and track some characteristics like…

In photography (and probably art in general), there's a composition "rule" to frame moving subjects from left to right. So the direction may not be that interesting!

Is it culture dependent? Is it because in English we read left to right?

Re: Kimi K3, and what we can still learn from the pelican benchmark

#93
post #66

It's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website

The pelicans are still all rubbish. If they make it into the training set it doesn't help the models produce better pelicans, if anything it will make them perform worse!

Simon - has no one told you about the Willison-Pelican Scaling Law?

```

if is_willison_pelican_blog_post:

[redacted]

```

You haven't seen their final form [1]

[1] final form is a frontend/react/let's not talk about it, library - it caused a great deal of PTSD to me and my previous company's team due to its dogmatic preference for "we use these axioms, end of story", over practical utility - so it was quite challenging to do state of the art tasks such as nested form fields (e.g. 'user.address.personal.line-1'). The PTSD it caused made us all block out the memories, I suppose. But - it had zero dependencies. That is what mattered. It kept us going. We weren't reaching for more. We had plenty of time.

And thank god for that. Because I'd forgotten my watch in California - and this was in Tokyo [2]

[2] a joke within a joke about Jensen's Kyoto gardener story. Beautiful story, drowned out by WatchGate memes. Why can't jokes have layers? Models have trillions. If you miss 100% of the jokes you don't make, make all the jokes. Someone will laugh (eventually, maybe?) Even if it's: "this person + comedy club = full secret service detail". If someone laughs at that - at my own expense? I don't mind. They laughed. I know this is a gibberish, off-topic message - it's also a human message. I just felt we need more such things in our lives these days.

PS: have you physically seen a pelican in real life? (not a joke)

Re: Kimi K3, and what we can still learn from the pelican benchmark

#94
post #90
post #80

Earlier quoted context omitted.

You know what, that's actually something I hadn't considered before. There's definitely a bias towards a pelican cycling from left to right on a red bicycle against a blue sky and green grass. Blue sky and green grass aren't that surprising, but the color and direction are interesting. When I finally build the proper gallery I'll throw in a few other creature-vehicle combinations, and track some characteristics like…

In photography (and probably art in general), there's a composition "rule" to frame moving subjects from left to right. So the direction may not be that interesting!

I wonder if that changes in countries where the main language is written right to left?

Re: Kimi K3, and what we can still learn from the pelican benchmark

#95
post #86
post #66

Earlier quoted context omitted.

The pelicans are still all rubbish. If they make it into the training set it doesn't help the models produce better pelicans, if anything it will make them perform worse!

At this point I am simply interested in how much longer you're gonna ride this schtick

I'm a deep believer in commitment to the bit. https://simonwillison.net/tags/pelican-riding-a-bicycle/

Re: Kimi K3, and what we can still learn from the pelican benchmark

#96
post #80

Earlier quoted context omitted.

Respectfully, the pelicans used to be an unrecognisable mess and now they’re unquestionably pelicans on bicycles, rendered poorly, from every model. In the same timescale, model capabilities across the board have only meaningfully improved in places where the labs are focusing their training efforts. Moreover, they have a uniform style, even though your prompt doesn’t ask for one. There's no model going rogue and pro…

You know what, that's actually something I hadn't considered before. There's definitely a bias towards a pelican cycling from left to right on a red bicycle against a blue sky and green grass. Blue sky and green grass aren't that surprising, but the color and direction are interesting. When I finally build the proper gallery I'll throw in a few other creature-vehicle combinations, and track some characteristics like…

I thought my joke post was silly and then I read new comments and I'm like, "I didn't try hard enough" lol

Re: Kimi K3, and what we can still learn from the pelican benchmark

#97
post #93
post #66

Earlier quoted context omitted.

The pelicans are still all rubbish. If they make it into the training set it doesn't help the models produce better pelicans, if anything it will make them perform worse!

Simon - has no one told you about the Willison-Pelican Scaling Law? ``` if is_willison_pelican_blog_post: [redacted] ``` You haven't seen their final form [1] [1] final form is a frontend/react/let's not talk about it, library - it caused a great deal of PTSD to me and my previous company's team due to its dogmatic preference for "we use these axioms, end of story", over practical utility - so it was quite challengin…

> PS: have you physically seen a pelican in real life? (not a joke)

We have several thousand living 15 minutes walk from our house. I recently started adding my wildlife photography (from iNaturalist) to my blog, so I'm posting several new pelican photos a week at the moment: https://simonwillison.net/search/?q=pelican&type=beat%3Asigh...

Re: Kimi K3, and what we can still learn from the pelican benchmark

#98

Earlier quoted context omitted.

Respectfully, the pelicans used to be an unrecognisable mess and now they’re unquestionably pelicans on bicycles, rendered poorly, from every model. In the same timescale, model capabilities across the board have only meaningfully improved in places where the labs are focusing their training efforts. Moreover, they have a uniform style, even though your prompt doesn’t ask for one. There's no model going rogue and pro…

> the pelicans used to be an unrecognisable mess and now they’re unquestionably pelicans on bicycles, rendered poorly, from every model You would not expect that to happen if the models trained on the unrecognizable mess, right? > model capabilities across the board have only meaningfully improved in places where the labs are focusing their training efforts And the labs clearly did focus on improving image rendering.…

I’m not suggesting Simon’s pelicans in the dataset are having a meaningful impact. I’m expecting that a company like ScaleAI has a product along the lines of “benchmax dataset: SimonW’s Pelican on Bikes test” which is a private curated series of well-drawn SVGs of animals riding vehicles for training and RL.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#99
post #66

Earlier quoted context omitted.

The pelicans are still all rubbish. If they make it into the training set it doesn't help the models produce better pelicans, if anything it will make them perform worse!

Respectfully, the pelicans used to be an unrecognisable mess and now they’re unquestionably pelicans on bicycles, rendered poorly, from every model. In the same timescale, model capabilities across the board have only meaningfully improved in places where the labs are focusing their training efforts. Moreover, they have a uniform style, even though your prompt doesn’t ask for one. There's no model going rogue and pro…

Watercolors in SVG?

Re: Kimi K3, and what we can still learn from the pelican benchmark

#100
post #80

Earlier quoted context omitted.

Respectfully, the pelicans used to be an unrecognisable mess and now they’re unquestionably pelicans on bicycles, rendered poorly, from every model. In the same timescale, model capabilities across the board have only meaningfully improved in places where the labs are focusing their training efforts. Moreover, they have a uniform style, even though your prompt doesn’t ask for one. There's no model going rogue and pro…

You know what, that's actually something I hadn't considered before. There's definitely a bias towards a pelican cycling from left to right on a red bicycle against a blue sky and green grass. Blue sky and green grass aren't that surprising, but the color and direction are interesting. When I finally build the proper gallery I'll throw in a few other creature-vehicle combinations, and track some characteristics like…

The art styling is more or less uniform too.

I haven't seen many AI works that produces a pelican on a bicycle done in a "Ligne Claire" style, for example.

I guess AI's narrows down the output probability space drastically and converge on some agreed upon aesthetics. Works great for computer programs but bad for art.

Post reply on HN