Live data from Hacker News

Kimi K3, and what we can still learn from the pelican benchmark

simonwillison.net

211–220 of 245 posts

Re: Kimi K3, and what we can still learn from the pelican benchmark

#213
post #201

Earlier quoted context omitted.

It's got nothing to do with what most people actually do when they're working.. AI companies claim their products are generalists though, and that they can do a good job on anything you give them, so you can't say what people will be doing with it. "Generate an SVG of an bird on a bicycle" is a corner case certainly but if a candidate interviewing for a role claims they can handle the corner cases then it's totally f…

I think they're less and less advertised as true generalists these days, as they pivot to profits that obviously lie (for the time being) first and foremost in agentic coding. It's no longer unusual to see regressions in terms of more stiff prose due to the strong tuning towards coding, or how they structure their response. And prose is a LLM's home turf! Instead, progress in agentic coding capability is usually the…

Anecdotally, GPT-3 was super good at creative writing. It didn’t have any of the typical LLM giveaways. It would write super weird, interesting stuff. Especially if fine-tuned on a specific author. Of course it would occasionally descend into saying the same thing over and over. But IMO none of the current models come close!

Re: Kimi K3, and what we can still learn from the pelican benchmark

#214
post #33

Earlier quoted context omitted.

Pelicans and bikes can be in the training set without them training for this specific benchmark.

Yes and that would improve its ability to draw SVGs of pelicans on bikes, no?

> Yes and that would improve its ability to draw SVGs of pelicans on bikes, no?

I would think the opposite because unless people have been hand drawing these with high quality, the training would be on much crappier versions that old AIs have done.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#215
post #86
post #66

Earlier quoted context omitted.

The pelicans are still all rubbish. If they make it into the training set it doesn't help the models produce better pelicans, if anything it will make them perform worse!

At this point I am simply interested in how much longer you're gonna ride this schtick

as long as he gets paid for that

Re: Kimi K3, and what we can still learn from the pelican benchmark

#216
post #209

I am not a fan of this benchmark, nor the interpretation of Simon's. Can you draw a pelican riding a bike, and that would pass with flying colors if ranked by a diverse set of human judges? If not, you have your answer r.e. test credibility.

That's the joke.

[dead]

Re: Kimi K3, and what we can still learn from the pelican benchmark

#217
post #147

Earlier quoted context omitted.

I think the user should be banned. It’s insane spam

I didn't submit this story. If you look at https://news.ycombinator.com/from?site=simonwillison.net you'll see that I submitted just one out of the last thirty articles from my site that were submitted to Hacker News - and the one I submitted failed to gain any votes.

lots of websites have their posts shadowbanned because of excessive spam. The amount of people that believes your blog should be in that list is growing.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#218
post #72

Don't see why we have to have this spammed every model release when Fable class models perform the same as Opus on basic tasks like these.

What spam? It’s one article. You can skip it

one article? more like 30 comments with a set of links to his blog.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#219

we can learn nothing from it apart from the large troll community that is HN that wants to do the same boring spiel every time a new model drops

don't blame the community for the work of one hustler and a permissive (just in this case) moderation.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#220

Imagine shilling some CLI tools no one uses in this post.

Lighten up. You’re reading a personal blog and complaining about an open source personal project he runs and distributes for free. He’s allowed to talk about his personal work on his personal blog. Especially considering the cli utility he talks about is directly related to the post. Imagine complaining about someone generating valuable content for free and not packaging it to your personal tastes.

> Imagine complaining about someone generating valuable content for free and not packaging it to your personal tastes.

We complain about spammers all the time, what's wrong with that?

Post reply on HN