Is there a gallery of all pelicans generated by simon over time?
Kimi K3, and what we can still learn from the pelican benchmark
41–50 of 245 posts
Re: Kimi K3, and what we can still learn from the pelican benchmark
#42It's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website
Did you read the post? It's not even that long. He explicitly mentions this...
Re: Kimi K3, and what we can still learn from the pelican benchmark
#43Re: Kimi K3, and what we can still learn from the pelican benchmark
#44Re: Kimi K3, and what we can still learn from the pelican benchmark
#45Re: Kimi K3, and what we can still learn from the pelican benchmark
#46K3 is as expensive as Sonnet, not great at writing English, is handing IP back to the Chinese, and once open source will be difficult to run at scale without the compute that OpenAI and Anthropic have largely grabbed. Sorry, how again is this the end of the frontier labs?
You mean the scale that AWS provides with Bedrock?
Re: Kimi K3, and what we can still learn from the pelican benchmark
#47Earlier quoted context omitted.
Yes and that would improve its ability to draw SVGs of pelicans on bikes, no?
and that is bad because ?
Re: Kimi K3, and what we can still learn from the pelican benchmark
#48It's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website
Did you read the post? It's not even that long. He explicitly mentions this...
Re: Kimi K3, and what we can still learn from the pelican benchmark
#49It's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website
Imagine if we applied this train of logic to humans. "That artist saw a pelican at the beach once!" [cue the outrage] "He's not a real artist, he's a cheater and produces nothing original!"
Re: Kimi K3, and what we can still learn from the pelican benchmark
#50Earlier quoted context omitted.
and that is bad because ?
the nature of the test was to see if the models can effectively compose an image of a novel concept outside the training set. If they are trained on it, it ceases to be an interesting test to some extent.