Live data from Hacker News

Kimi K3, and what we can still learn from the pelican benchmark

simonwillison.net

61–70 of 245 posts

Re: Kimi K3, and what we can still learn from the pelican benchmark

#61
post #33

Earlier quoted context omitted.

Pelicans and bikes can be in the training set without them training for this specific benchmark.

Yes and that would improve its ability to draw SVGs of pelicans on bikes, no?

Would it? Tongue in cheek.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#62

It’s not bad kind of expensive for 25c but if the prompt is rendered cost is much better.

I wonder what the non-subsidized cost is. Add in the electricity and water too.

We may be boiling the oceans but at least we are finally getting some good SVGs of pelicans on bicycles.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#63
Like Simon concludes the article, the main use of this isn't to say which model is "better", but to try and poke at the model to sort out things like quality vs cost vs speed.

So I put together a quick comparison of the last couple iterations of Opus, Fable and now Kimi.

Kimi is cheapest by 5x but also slowest by 2x

https://9gpyw4uxr2.evvl.io/

Re: Kimi K3, and what we can still learn from the pelican benchmark

#64
3T is impressive, but parameter count seems to be less important than I thought.

GLM is half the size of DeepSeek but costs four times as much, and beats it on every benchmark.

I'm not an expert on this stuff but it seems to be the attention mechanism. DeepSeek were bragging about how cheap they made it. But if you cut costs on attention you get worse results with way more parameters.

If I had to guess it seems to be the difference between memory (params) and intelligence (attention density). I think you need both.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#66

It's incredible Simon still believes pelicans on bikes aren't part of the training set, despite hundreds of them on blogs, forums, and Github. Stuff we put in our company blog shows up known by LLMs 6 months later, and we have 1000x less traffic than Simon's own website

The pelicans are still all rubbish. If they make it into the training set it doesn't help the models produce better pelicans, if anything it will make them perform worse!

Re: Kimi K3, and what we can still learn from the pelican benchmark

#67
post #45

I wonder how the Chinese labs are training a 3 trillion parameter model on what has to be vastly smaller compute resources. If the U.S. compute advantage is persistent, it's hard to imagine that Chinese labs will be able to keep pace forever, as a matter of physics, but... so far they seem to be doing just fine.

It's not like same parameter count models are identical, so that doesn't appear to be an indicator for quality, or even compute requirements?

There seems to be more to producing a better model than brute forcing parameter count after all.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#68
post #60
post #48

Earlier quoted context omitted.

Clearly not. There's a subset of HN users who rush to post this same thing every single time.

Maybe it gets posted every time because besides a personal believe by the person popularising this "benchmark", there is no reason to assume that certain labs aren't intentionally training to game this and every other lab at least unintentionally gets improvements for this specific combination of animal and action because the internet is full of both good and bad examples, often ranked, which does inevitably become t…

> I have shared examples of certain models by certain labs doing far better on the pelican cycling vs other, similar prompts

Please share those again!

One of the things I'm most looking forward to is a lab producing a model that creates a really great pelican riding a bicycle and then a terrible sloth riding a skateboard (or whatever).

I've not seen that myself yet.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#70
post #10
post #3

Another day, another model and another pelican :-) I can't help but wonder where is the trend going? What will we have in five years? Maybe it will all have puttered out, and we will have moved to the next thing? Or maybe the prompt then will be "make a pelican ride a bicycle", and out will come the genetic code for a giant pelican with extremities suitable for a handle bar and pedals, and an inborn affinity to ride…

You are thinking too hard on this. This entire "benchmark" is a performative joke for attention that only works on HN. > What will we have in five years? Maybe it will all have puttered out, and we will have moved to the next thing? We will just have more of the same.

> This entire "benchmark" is a performative joke for attention that only works on HN.

I take exception to that! It's a performative joke for attention that works far more widely than just Hacker News.

Post reply on HN