Live data from Hacker News

Kimi K3, and what we can still learn from the pelican benchmark

simonwillison.net

41–50 of 245 posts

Re: Kimi K3, and what we can still learn from the pelican benchmark

#42

It's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website

Did you read the post? It's not even that long. He explicitly mentions this...

Are they responding to: “I’m still not convinced that labs are training for the benchmark—if they were, I’d expect much better results.”

Re: Kimi K3, and what we can still learn from the pelican benchmark

#43
post #33

Earlier quoted context omitted.

Pelicans and bikes can be in the training set without them training for this specific benchmark.

Yes and that would improve its ability to draw SVGs of pelicans on bikes, no?

and that is bad because ?

Re: Kimi K3, and what we can still learn from the pelican benchmark

#45
I wonder how the Chinese labs are training a 3 trillion parameter model on what has to be vastly smaller compute resources. If the U.S. compute advantage is persistent, it's hard to imagine that Chinese labs will be able to keep pace forever, as a matter of physics, but... so far they seem to be doing just fine.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#46
post #28
post #23

K3 is as expensive as Sonnet, not great at writing English, is handing IP back to the Chinese, and once open source will be difficult to run at scale without the compute that OpenAI and Anthropic have largely grabbed. Sorry, how again is this the end of the frontier labs?

You mean the scale that AWS provides with Bedrock?

Bedrock needs to actually update their chinese models to the newest versions for this to matter.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#47
post #43
post #33

Earlier quoted context omitted.

Yes and that would improve its ability to draw SVGs of pelicans on bikes, no?

and that is bad because ?

the nature of the test was to see if the models can effectively compose an image of a novel concept outside the training set. If they are trained on it, it ceases to be an interesting test to some extent.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#48

It's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website

Did you read the post? It's not even that long. He explicitly mentions this...

Clearly not. There's a subset of HN users who rush to post this same thing every single time.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#49

It's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website

Imagine if we applied this train of logic to humans. "That artist saw a pelican at the beach once!" [cue the outrage] "He's not a real artist, he's a cheater and produces nothing original!"

This is a sight-reading test. If a musician practices a piece for thousands of hours, it would no longer be an effective sight reading / creativity test. The purpose of the test was to see how models would compose something novel requiring the ability to compose orthogonal, normally unrelated, components into a coherent image.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#50
post #43

Earlier quoted context omitted.

and that is bad because ?

the nature of the test was to see if the models can effectively compose an image of a novel concept outside the training set. If they are trained on it, it ceases to be an interesting test to some extent.

it's still interesting because there's no pelican-on-bike model, and if you're training a model well enough, then it should be obvious when a model has reached "AGI" or whatever.
Post reply on HN