Live data from Hacker News

Kimi K3, and what we can still learn from the pelican benchmark

simonwillison.net

31–40 of 245 posts

Re: Kimi K3, and what we can still learn from the pelican benchmark

#31
post #15

> This is expensive—the pelican cost 25 cents! Engineers get unbelievably silly about evaluating costs of things. "The tokens are so expensive!" Oh my sweet child, how much would even the least capable human effort cost? This is what the executives properly understand that the programmers don't.

Would anyone pay a human to create an SVG of a pelican riding a bike?

Well, no, not now they won’t.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#33

It's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website

Pelicans and bikes can be in the training set without them training for this specific benchmark.

Yes and that would improve its ability to draw SVGs of pelicans on bikes, no?

Re: Kimi K3, and what we can still learn from the pelican benchmark

#34

It's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website

Imagine if we applied this train of logic to humans.

"That artist saw a pelican at the beach once!" [cue the outrage] "He's not a real artist, he's a cheater and produces nothing original!"

Re: Kimi K3, and what we can still learn from the pelican benchmark

#35
post #24
post #10

Earlier quoted context omitted.

You are thinking too hard on this. This entire "benchmark" is a performative joke for attention that only works on HN. > What will we have in five years? Maybe it will all have puttered out, and we will have moved to the next thing? We will just have more of the same.

You say it's performative joke, but it all depends what you're using model for. So far the rule has been quite straightforward, better models consistently renders pelican in higher quality, I've yet to see an exception. It is also a good enough (for me at least) test for "taste" the model has.

> better models consistently renders pelican in higher quality The article literally avoid making this argument and gives counterexamples to this statement.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#36

It's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website

Imagine if we applied this train of logic to humans. "That artist saw a pelican at the beach once!" [cue the outrage] "He's not a real artist, he's a cheater and produces nothing original!"

Except, of course, LLMs are not humans, and they do not learn or "reason" in a way which even remotely resembles humans.

Plus obviously humans can still overfit to a specific style of test.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#37

It's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website

Did you read the post? It's not even that long. He explicitly mentions this...

Re: Kimi K3, and what we can still learn from the pelican benchmark

#38
If anyone wants to try SVG generation from different models, I made this: https://codeinput.com/svg (here is an older generation: https://codeinput.com/s/5KEGl1e3rB3)

You still need an OpenRouter API Key and be careful this can burn quite a bit of money.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#39

Imagine shilling some CLI tools no one uses in this post.

Lighten up.

You’re reading a personal blog and complaining about an open source personal project he runs and distributes for free. He’s allowed to talk about his personal work on his personal blog. Especially considering the cli utility he talks about is directly related to the post.

Imagine complaining about someone generating valuable content for free and not packaging it to your personal tastes.

Post reply on HN