Live data from Hacker News

Kimi K3, and what we can still learn from the pelican benchmark

simonwillison.net

201–210 of 245 posts

Re: Kimi K3, and what we can still learn from the pelican benchmark

#201

The pelican benchmark is exactly what's wrong with hiring in technology. It's got nothing to do with what most people actually do when they're working - just like most job interviews which ask you to draw a pelican as their way of assessing you.

It's got nothing to do with what most people actually do when they're working.. AI companies claim their products are generalists though, and that they can do a good job on anything you give them, so you can't say what people will be doing with it. "Generate an SVG of an bird on a bicycle" is a corner case certainly but if a candidate interviewing for a role claims they can handle the corner cases then it's totally f…

I think they're less and less advertised as true generalists these days, as they pivot to profits that obviously lie (for the time being) first and foremost in agentic coding. It's no longer unusual to see regressions in terms of more stiff prose due to the strong tuning towards coding, or how they structure their response. And prose is a LLM's home turf! Instead, progress in agentic coding capability is usually the headline feature, the headline benchmark, etc etc. At least looking at Anthropic, Google, OpenAI. There are of course other LLM's.

So then add a dash of cybersecurity and medical use and that's basically it. No "closer to AGI" advertising. I'd say the 2026 development has in fact been the opposite; optimizing AI for niches where there is most potential for profits and that your description died in circa GPT-5 era.

In fact, this problem (for this test) is also stated by the pelican test author:

"The biggest limitation of the pelican is that it doesn’t touch at all on the thing that matters most for today’s model: agentic tool calling and the ability to operate tools reliably as conversations grow in length.

So don’t go using pelicans to compare models!"

Re: Kimi K3, and what we can still learn from the pelican benchmark

#203
post #201

Earlier quoted context omitted.

It's got nothing to do with what most people actually do when they're working.. AI companies claim their products are generalists though, and that they can do a good job on anything you give them, so you can't say what people will be doing with it. "Generate an SVG of an bird on a bicycle" is a corner case certainly but if a candidate interviewing for a role claims they can handle the corner cases then it's totally f…

I think they're less and less advertised as true generalists these days, as they pivot to profits that obviously lie (for the time being) first and foremost in agentic coding. It's no longer unusual to see regressions in terms of more stiff prose due to the strong tuning towards coding, or how they structure their response. And prose is a LLM's home turf! Instead, progress in agentic coding capability is usually the…

LLMs are, fundamentally, generalist AIs. Marketing or no marketing - it's just what they are. How they're trained, how they perform, what they're best at.

Empirically, they have something very much alike to the human "g factor" - a shared pool of "general intelligence" that all tasks benefit from.

When a "make it bigger, train it harder" upgrade like Kimi K3 or Mythos 5 drops, the performance rises on every metric. Not just the "headline benchmarks" like Mythos and coding/cybersecurity, but also things like literary analysis - which has nearly zero economic value, and isn't commonly post-trained or benchmarked for. And companies keep encountering things like "our carefully trained specialist model with lots of in-domain training on expensive closed datasets just got leapfrogged on our internal benchmarks by a next gen off the shelf generalist".

You can go hard on benchmarkmaxxing post-training, and you can burn millions of GPU-hours on coding RLVR. But, by the very nature of LLMs, a lot of the performance gains in flagship models are broad and domain-inspecific.

"Stiff prose" is more of a "style" thing than a "capability" thing. No one cares about how good an AI is at things like long form creative writing, because that's the opposite of a profitable field. All of LLM behavior is routed through text, so it's very easy to perturb "writing style" by some training elsewhere. Regression evaluation is hard. And the writing-specific post-training LLMs get is usually just cheap RLAF, with all the usual RLAF degeneracy.

Thus, we get the "default styles" that suck from a "creative writing" standpoint. A lot of that is just "what sounded good to the previous generation of LLMs" - and, unlike human readers, LLM evaluators don't get bored from seeing the same cliches repeated 9000 times across 9000 different instances of generated text. Humans tend to update over time from "this sound cool and punchy" to "this is generic AI slop", but RLAF evaluators stay at step 1. What little human-guided optimization this gets is aimed at "copywriting, marketing blurbs, punchy short-form" - and it shows.

You can do a lot there with some aggressive prompting, but the default writing styles suck, and I frankly don't expect that to change soon. No one cares enough to change it.

Pelicans? Used to be a decent proxy for "general model capabilities that no one would benchmaxx for" - a way to probe for that elusive "LLM g factor". Now that it's a known metric, it's very gameable. But it was pretty solid while it was novel and obscure.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#205
post #80

Earlier quoted context omitted.

You know what, that's actually something I hadn't considered before. There's definitely a bias towards a pelican cycling from left to right on a red bicycle against a blue sky and green grass. Blue sky and green grass aren't that surprising, but the color and direction are interesting. When I finally build the proper gallery I'll throw in a few other creature-vehicle combinations, and track some characteristics like…

It'd be hard to fully compare, but I think a truly random "creature-vehicle" along side the pelican test would catch who's gaming and who's not. I'd also enjoy the absurdism of "Herring on a pogostick"

The models are already brilliant at that. My own todo app generates 128x128 pixel art icons for my todo items. They are mind blowingly creative and funny.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#206
post #42

Earlier quoted context omitted.

Did you read the post? It's not even that long. He explicitly mentions this...

Are they responding to: “I’m still not convinced that labs are training for the benchmark—if they were, I’d expect much better results.”

In my reading, "training for the benchmark" is very, very different from "this benchmark is in the training data".

Re: Kimi K3, and what we can still learn from the pelican benchmark

#207

The pelican benchmark is exactly what's wrong with hiring in technology. It's got nothing to do with what most people actually do when they're working - just like most job interviews which ask you to draw a pelican as their way of assessing you.

It is a simple prompt that packs a lot:

* operates an absurd prompt

* involves SVG coding knowledge, generates a source code artifact

* involves world knowledge (what is a pelican? What is a bicycle? What does each do?” How are each constructed?”)

* when rendered, the coding artifact expresses an image that makes sense to us perceptually, including color and spatial relationships

* different models and settings have different output so it can be used as an evaluation scheme

That said I wouldn’t choose a model based on this! Just like some brain teaser shouldn’t determine employment eligibility.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#208

Earlier quoted context omitted.

Or they just don't actually have any compute access restrictions of significance? Chinese companies can just go use those GPUs in neighboring countries that aren't export-restricted, like Malaysia. Like ByteDance openly did: https://www.tomshardware.com/pc-components/gpus/chinas-byted... and Tencent is rumored to have done via Japan: https://wccftech.com/china-tencent-gains-access-to-nvidia-bl... And that's not even…

Even ignoring chip export ban, Chinese companies have way less funding than American counterparts, maybe 1 or 2 orders of magnitude less depending on which company you look at. Deepseek’s recent big funding round being “only” a couple billion $ at $50B valuation, for example. Bytedance and Tencent are tech giants for sure, nonetheless they’re not Google kind of giant.

> Bytedance and Tencent are tech giants for sure, nonetheless they’re not Google kind of giant.

$186 billion and $105 billion revenue in 2025 respectively vs. $402 billion? Yes, Google is larger, but they're all in that same ballpark?

ByteDance's 2025 net income isn't that different from Anthropic's Series H funding even ($50bn vs $65bn respectively).

But this is all also ignoring how much of China is state owned (25% of the GDP!), so the available resource pool is dramatically larger than it would appear depending on what the government decides is important

Re: Kimi K3, and what we can still learn from the pelican benchmark

#209

I am not a fan of this benchmark, nor the interpretation of Simon's. Can you draw a pelican riding a bike, and that would pass with flying colors if ranked by a diverse set of human judges? If not, you have your answer r.e. test credibility.

That's the joke.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#210

The pelican benchmark is exactly what's wrong with hiring in technology. It's got nothing to do with what most people actually do when they're working - just like most job interviews which ask you to draw a pelican as their way of assessing you.

It's got nothing to do with what most people actually do when they're working.. AI companies claim their products are generalists though, and that they can do a good job on anything you give them, so you can't say what people will be doing with it. "Generate an SVG of an bird on a bicycle" is a corner case certainly but if a candidate interviewing for a role claims they can handle the corner cases then it's totally f…

Just as even a counterpoint to this, I have asked the LLMs to attempt to generate SVG icons for websites. Even though I have requested things much simpler than a Pelican, they have all tended to do quite poorly in my examples.

Because of this, I presume the Pelican has been in the training data for at least a year+.

The models are very useful, I am afraid they have fundamental limitations though generalizing (it is just hard to evaluate effectively). So it will just be whack-a-mole "can your model do X", and there will always be a new X.

Post reply on HN