Kimi K3, and what we can still learn from the pelican benchmark
51–60 of 245 posts
Re: Kimi K3, and what we can still learn from the pelican benchmark
#52Re: Kimi K3, and what we can still learn from the pelican benchmark
#53I would be surprised if pelican svgs are not part of the training corpus rn
If that were the case then it'd do a way better job. Think experienced artist level.
what they do have are many different pelicans and people helpfully rating them in the comments.
Re: Kimi K3, and what we can still learn from the pelican benchmark
#54Earlier quoted context omitted.
and that is bad because ?
the nature of the test was to see if the models can effectively compose an image of a novel concept outside the training set. If they are trained on it, it ceases to be an interesting test to some extent.
Re: Kimi K3, and what we can still learn from the pelican benchmark
#55It's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website
Simon has stated a few times that he knows it’s possible that pelicans could be in the training sets. He also has other tests he doesn’t share publicly. He’s just a fan of pelicans.
Re: Kimi K3, and what we can still learn from the pelican benchmark
#56It's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website
Did you read the post? It's not even that long. He explicitly mentions this...
Did you read either the post or the comment it was referencing?
On the note of training on SVGs, I have seen some labs models outperform when prompted for SVGs of certain animal and action combinations (pelican on bike, panda eating burger, etc.) compared to other similarly outlandish prompts for SVG output that are not part of widely reported benchmarks, even shared evidence one of the last times this came up on here.
[0] ... incredible Simon still believes ...
[1] I’m still not convinced that labs ....
Re: Kimi K3, and what we can still learn from the pelican benchmark
#57It's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website
Re: Kimi K3, and what we can still learn from the pelican benchmark
#58It's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website
Imagine if we applied this train of logic to humans. "That artist saw a pelican at the beach once!" [cue the outrage] "He's not a real artist, he's a cheater and produces nothing original!"
Re: Kimi K3, and what we can still learn from the pelican benchmark
#59Do any of the vision models render the SVG and look at the result. Perhaps more importantly can they do that during reinforcement training. Learning how to critically analyse the appearance of what it generates would be quite useful. Manually feeding images back to models has been hilariously bad in the past which suggests that relating something it sees to something it wrote is not an ability it is very good at.
Re: Kimi K3, and what we can still learn from the pelican benchmark
#60Earlier quoted context omitted.
Did you read the post? It's not even that long. He explicitly mentions this...
Clearly not. There's a subset of HN users who rush to post this same thing every single time.
I have shared examples of certain models by certain labs doing far better on the pelican cycling vs other, similar prompts. Just operating on a feeling that labs don't optimise for this (as mentioned, even if they don't training data is filled with these) is not solid enough that criticism shouldn't be leveraged when it comes up.