Live data from Hacker News

Kimi K3, and what we can still learn from the pelican benchmark

simonwillison.net

101–110 of 245 posts

Re: Kimi K3, and what we can still learn from the pelican benchmark

#101

Earlier quoted context omitted.

> the pelicans used to be an unrecognisable mess and now they’re unquestionably pelicans on bicycles, rendered poorly, from every model You would not expect that to happen if the models trained on the unrecognizable mess, right? > model capabilities across the board have only meaningfully improved in places where the labs are focusing their training efforts And the labs clearly did focus on improving image rendering.…

I’m not suggesting Simon’s pelicans in the dataset are having a meaningful impact. I’m expecting that a company like ScaleAI has a product along the lines of “benchmax dataset: SimonW’s Pelican on Bikes test” which is a private curated series of well-drawn SVGs of animals riding vehicles for training and RL.

If such a product existed I'm reasonably confident someone would have tipped me off by now, NDAs be damned.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#102
post #64

3T is impressive, but parameter count seems to be less important than I thought. GLM is half the size of DeepSeek but costs four times as much, and beats it on every benchmark. I'm not an expert on this stuff but it seems to be the attention mechanism. DeepSeek were bragging about how cheap they made it. But if you cut costs on attention you get worse results with way more parameters. If I had to guess it seems to be…

After MoE entered the mix, raw parameter count is less useful a measure.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#103
post #92
post #90

Earlier quoted context omitted.

In photography (and probably art in general), there's a composition "rule" to frame moving subjects from left to right. So the direction may not be that interesting!

Is it culture dependent? Is it because in English we read left to right?

There was a glorious moment when I thought that the Chinese models were more likely to produce right-to-left cycling pelicans, but sadly that trend didn't seem to hold up.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#104
post #45

I wonder how the Chinese labs are training a 3 trillion parameter model on what has to be vastly smaller compute resources. If the U.S. compute advantage is persistent, it's hard to imagine that Chinese labs will be able to keep pace forever, as a matter of physics, but... so far they seem to be doing just fine.

It's not like same parameter count models are identical, so that doesn't appear to be an indicator for quality, or even compute requirements? There seems to be more to producing a better model than brute forcing parameter count after all.

Training and serving large models does require increasingly more compute, though. (The Chinese labs have clearly found some massive optimizations, but my point was that you'd think at some point even those optimizations wouldn't be enough to keep up with exponentially increasing model sizes.)

Re: Kimi K3, and what we can still learn from the pelican benchmark

#105
post #64

3T is impressive, but parameter count seems to be less important than I thought. GLM is half the size of DeepSeek but costs four times as much, and beats it on every benchmark. I'm not an expert on this stuff but it seems to be the attention mechanism. DeepSeek were bragging about how cheap they made it. But if you cut costs on attention you get worse results with way more parameters. If I had to guess it seems to be…

Or, GLM 5.2 simply had more time in the RL oven.

Deepseek V4 Flash, the 284B model, is roughly equivalent to launch GLM 5, the 744B [sic] model.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#106
post #103
post #92

Earlier quoted context omitted.

Is it culture dependent? Is it because in English we read left to right?

There was a glorious moment when I thought that the Chinese models were more likely to produce right-to-left cycling pelicans, but sadly that trend didn't seem to hold up.

Chinese is also written left to right

Re: Kimi K3, and what we can still learn from the pelican benchmark

#107
post #21

Imagine what amazing SVG generators we could have if Simon had randomized the target image from the start (and companies wouldn't just overfit on pelicans).

I think a pelican riding a bike is fairly random. (https://xkcd.com/221/)

Re: Kimi K3, and what we can still learn from the pelican benchmark

#108
post #103
post #92

Earlier quoted context omitted.

Is it culture dependent? Is it because in English we read left to right?

There was a glorious moment when I thought that the Chinese models were more likely to produce right-to-left cycling pelicans, but sadly that trend didn't seem to hold up.

For almost the last 70 years, Chinese has been left to right.

Before that it was vertical (although the ordering of the columns was right to left).

Re: Kimi K3, and what we can still learn from the pelican benchmark

#109
post #82
post #56

Earlier quoted context omitted.

Respectfully, did you? The comment was specific to doubting the believe simonw has that labs are not training [0] specifically for this task, which is exactly what simonw wrote in the post [1], that it is a believe of his that they don't. He did not mention any kind of evidence or any piece of information that would indicate that the commenter didn't read the blog post. Did you read either the post or the comment it…

I'll note there's a difference between "pelicans on bikes aren't part of the training set" and "I’m still not convinced that labs are training for the benchmark". I'm sure all sorts of crap pelican riding bicycle SVGs have ended up in the huge crawls of data that the labs feed into their pre-training steps. What I'm questioning here is that there are labs who have sat down and deliberately tested and tweaked the perf…

> What I'm questioning here is that there are labs who have sat down and deliberately tested and tweaked the performance for this particular task, independent of general model improvements.

Given the massive delta easily reproducible with some models, is it really doubtful that certain labs have not: https://news.ycombinator.com/item?id=48951229

We are going from pretty good pelican to jumbled mess with a similarly silly, but different prompt across multiple models from multiple labs, both Western and Eastern, both Open Weight and Closed.

Post reply on HN