Live data from Hacker News

Kimi K3, and what we can still learn from the pelican benchmark

simonwillison.net

21–30 of 245 posts

Re: Kimi K3, and what we can still learn from the pelican benchmark

#23
K3 is as expensive as Sonnet, not great at writing English, is handing IP back to the Chinese, and once open source will be difficult to run at scale without the compute that OpenAI and Anthropic have largely grabbed.

Sorry, how again is this the end of the frontier labs?

Re: Kimi K3, and what we can still learn from the pelican benchmark

#24
post #10
post #3

Another day, another model and another pelican :-) I can't help but wonder where is the trend going? What will we have in five years? Maybe it will all have puttered out, and we will have moved to the next thing? Or maybe the prompt then will be "make a pelican ride a bicycle", and out will come the genetic code for a giant pelican with extremities suitable for a handle bar and pedals, and an inborn affinity to ride…

You are thinking too hard on this. This entire "benchmark" is a performative joke for attention that only works on HN. > What will we have in five years? Maybe it will all have puttered out, and we will have moved to the next thing? We will just have more of the same.

You say it's performative joke, but it all depends what you're using model for. So far the rule has been quite straightforward, better models consistently renders pelican in higher quality, I've yet to see an exception. It is also a good enough (for me at least) test for "taste" the model has.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#25
post #23

K3 is as expensive as Sonnet, not great at writing English, is handing IP back to the Chinese, and once open source will be difficult to run at scale without the compute that OpenAI and Anthropic have largely grabbed. Sorry, how again is this the end of the frontier labs?

According to some benchmarks has the coding capability of Opus at the price of Sonnet, supposedly will be open weights and is not subject to random trade wars with allied states.

Competition is always good.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#27

It's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website

More to it, the actual bloody companies are using them as a reference. Maybe it’s a 3d version, not an svg - but it clearly shows they’re on the radar of these companies.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#28
post #23

K3 is as expensive as Sonnet, not great at writing English, is handing IP back to the Chinese, and once open source will be difficult to run at scale without the compute that OpenAI and Anthropic have largely grabbed. Sorry, how again is this the end of the frontier labs?

You mean the scale that AWS provides with Bedrock?

Re: Kimi K3, and what we can still learn from the pelican benchmark

#29

It's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website

Simon has stated a few times that he knows it’s possible that pelicans could be in the training sets. He also has other tests he doesn’t share publicly. He’s just a fan of pelicans.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#30
Do any of the vision models render the SVG and look at the result.

Perhaps more importantly can they do that during reinforcement training. Learning how to critically analyse the appearance of what it generates would be quite useful.

Manually feeding images back to models has been hilariously bad in the past which suggests that relating something it sees to something it wrote is not an ability it is very good at.

Post reply on HN