Live data from Hacker News

Kimi K3, and what we can still learn from the pelican benchmark

simonwillison.net

141–150 of 245 posts

Re: Kimi K3, and what we can still learn from the pelican benchmark

#142
post #97
post #93

Earlier quoted context omitted.

Simon - has no one told you about the Willison-Pelican Scaling Law? ``` if is_willison_pelican_blog_post: [redacted] ``` You haven't seen their final form [1] [1] final form is a frontend/react/let's not talk about it, library - it caused a great deal of PTSD to me and my previous company's team due to its dogmatic preference for "we use these axioms, end of story", over practical utility - so it was quite challengin…

> PS: have you physically seen a pelican in real life? (not a joke) We have several thousand living 15 minutes walk from our house. I recently started adding my wildlife photography (from iNaturalist) to my blog, so I'm posting several new pelican photos a week at the moment: https://simonwillison.net/search/?q=pelican&type=beat%3Asigh...

Simon - thank you for not dismissing it (and surviving the text that came before the question).

I asked because I genuinely feel that the % of people working on some of the most important technology these days - things such as these 'strangely shaped tools' (to borrow from nearcyan) - large language models - the younger generation (folks in their early/mid 20s) - it is not unlikely that they have not physically seen the meatspace version of whatever digital correspondence of it that is being packed into latent space.

After all, why waste time going to the SF or Oakland zoo? One can just check Simon's latest pelican blog post and skip the zoo trip - the harnesses are waiting.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#144
post #99

Earlier quoted context omitted.

Respectfully, the pelicans used to be an unrecognisable mess and now they’re unquestionably pelicans on bicycles, rendered poorly, from every model. In the same timescale, model capabilities across the board have only meaningfully improved in places where the labs are focusing their training efforts. Moreover, they have a uniform style, even though your prompt doesn’t ask for one. There's no model going rogue and pro…

Watercolors in SVG?

Also, I'd assume the ideal output for an underspecified, generic prompt is the most expected, generic result. Not something that defaults off the rails with creative license.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#145
post #66

Earlier quoted context omitted.

The pelicans are still all rubbish. If they make it into the training set it doesn't help the models produce better pelicans, if anything it will make them perform worse!

Respectfully, the pelicans used to be an unrecognisable mess and now they’re unquestionably pelicans on bicycles, rendered poorly, from every model. In the same timescale, model capabilities across the board have only meaningfully improved in places where the labs are focusing their training efforts. Moreover, they have a uniform style, even though your prompt doesn’t ask for one. There's no model going rogue and pro…

being able to draw a picture of a pelican is really cool and it requires intelligence but i don't think it's a good measure of improving capabilities of these models nor AGI. we don't have to spend so much breath on it.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#146
post #82
post #56

Earlier quoted context omitted.

Respectfully, did you? The comment was specific to doubting the believe simonw has that labs are not training [0] specifically for this task, which is exactly what simonw wrote in the post [1], that it is a believe of his that they don't. He did not mention any kind of evidence or any piece of information that would indicate that the commenter didn't read the blog post. Did you read either the post or the comment it…

I'll note there's a difference between "pelicans on bikes aren't part of the training set" and "I’m still not convinced that labs are training for the benchmark". I'm sure all sorts of crap pelican riding bicycle SVGs have ended up in the huge crawls of data that the labs feed into their pre-training steps. What I'm questioning here is that there are labs who have sat down and deliberately tested and tweaked the perf…

Yes that's the obvious thing to do and why straightforward variants of known tests would also be treated as contaminated by anyone being even somewhat rigorous.

I don't know why the standard is is to be sure that it is happening versus it being a plausible risk of making the results useless.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#147

Don't see why we have to have this spammed every model release when Fable class models perform the same as Opus on basic tasks like these.

I think the user should be banned. It’s insane spam

I didn't submit this story.

If you look at https://news.ycombinator.com/from?site=simonwillison.net you'll see that I submitted just one out of the last thirty articles from my site that were submitted to Hacker News - and the one I submitted failed to gain any votes.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#149

It's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website

It's incredible you can't reason to see if pelican on a bike is a thing. It's not! This has been discussed to death. You can ask any model to generate anything. Generate an SVG of earthworm and a robin boxing. Guess what? The smarter the model the better the image, doesn't matter if it's a vision model or not. I rolled my eyes at this eval when I first saw it, then I tried various ridiculous things and noticed a very strong correlation. Things that are absolutely not in the training set.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#150
post #80

Earlier quoted context omitted.

Respectfully, the pelicans used to be an unrecognisable mess and now they’re unquestionably pelicans on bicycles, rendered poorly, from every model. In the same timescale, model capabilities across the board have only meaningfully improved in places where the labs are focusing their training efforts. Moreover, they have a uniform style, even though your prompt doesn’t ask for one. There's no model going rogue and pro…

You know what, that's actually something I hadn't considered before. There's definitely a bias towards a pelican cycling from left to right on a red bicycle against a blue sky and green grass. Blue sky and green grass aren't that surprising, but the color and direction are interesting. When I finally build the proper gallery I'll throw in a few other creature-vehicle combinations, and track some characteristics like…

There's a bias in the direction all things face. You can ask these models to generate a thing animal, car etc and you will notice that 90% of them will converge towards the same sort of results. If you ask for something rotating, 90% of them will rotate right and a few odd ones will rotate left.
Post reply on HN