Live data from Hacker News

Kimi K3, and what we can still learn from the pelican benchmark

simonwillison.net

221–230 of 245 posts

Re: Kimi K3, and what we can still learn from the pelican benchmark

#221
post #66

Earlier quoted context omitted.

The pelicans are still all rubbish. If they make it into the training set it doesn't help the models produce better pelicans, if anything it will make them perform worse!

Respectfully, the pelicans used to be an unrecognisable mess and now they’re unquestionably pelicans on bicycles, rendered poorly, from every model. In the same timescale, model capabilities across the board have only meaningfully improved in places where the labs are focusing their training efforts. Moreover, they have a uniform style, even though your prompt doesn’t ask for one. There's no model going rogue and pro…

> Moreover, they have a uniform style, even though your prompt doesn’t ask for one.

This shouldn’t really come as a surprise, particularly to anyone who’s used diffusion models. The same thing happens when you ask an LLM for a short story [1] without providing any specific details.

Even cranking up the temperature or top_p values is no panacea. The more generic your prompt, the more pedestrian the response.

[1] - https://news.ycombinator.com/item?id=42093394

Re: Kimi K3, and what we can still learn from the pelican benchmark

#222
post #140

> The biggest limitation of the pelican is that it doesn’t touch at all on the thing that matters most for today’s model: agentic tool calling and the ability to operate tools reliably as conversations grow in length. In all seriousness, I propose SWE-bench-adversarial-pelican-gen: it's like SWE-bench, but the harness gets interrupted every 5 turns/tool-calls and is asked to produce an SVG of an arbitrary animal befo…

Ask it to write a program that outputs SVGs of animals using human modes of transportation, then run the program with "pelican" and "bicycle" as inputs.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#223

Earlier quoted context omitted.

It's got nothing to do with what most people actually do when they're working.. AI companies claim their products are generalists though, and that they can do a good job on anything you give them, so you can't say what people will be doing with it. "Generate an SVG of an bird on a bicycle" is a corner case certainly but if a candidate interviewing for a role claims they can handle the corner cases then it's totally f…

Just as even a counterpoint to this, I have asked the LLMs to attempt to generate SVG icons for websites. Even though I have requested things much simpler than a Pelican, they have all tended to do quite poorly in my examples. Because of this, I presume the Pelican has been in the training data for at least a year+. The models are very useful, I am afraid they have fundamental limitations though generalizing (it is j…

Exactly this! I've tried to generate some really basic SVG icons (think fontawesome) with sota models (one generation back - so gpt 5.5) and _none_ have produced anything that I could use as-is and I've needed to fix stuff in the SVGs manually.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#224
post #147

Earlier quoted context omitted.

I didn't submit this story. If you look at https://news.ycombinator.com/from?site=simonwillison.net you'll see that I submitted just one out of the last thirty articles from my site that were submitted to Hacker News - and the one I submitted failed to gain any votes.

lots of websites have their posts shadowbanned because of excessive spam. The amount of people that believes your blog should be in that list is growing.

The amount of people who vote them up appears to be significantly higher.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#228

The pelican benchmark is exactly what's wrong with hiring in technology. It's got nothing to do with what most people actually do when they're working - just like most job interviews which ask you to draw a pelican as their way of assessing you.

Hi, experienced pelican here. No.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#229
> The new model is notable for the pricing: $3/million input tokens and $15/million output tokens, putting it at the same level as Anthropic’s Claude Sonnet series and making it the most expensive model released by a Chinese AI lab to date

Interesting to see that prices are converging to an "equilibrium price" regardless of being US or Chinese.

Re: Kimi K3, and what we can still learn from the pelican benchmark

#230

I read it. I did not get what we can still learn from this benchmark, and I still don’t understand what’s the point of it but sure.

Agree can we stop talking about this? Simon has had his time, but nothing interesting coming anymore. Last posts have been shots in the oven.
Post reply on HN