Live data from Hacker News

The last six months in LLMs, illustrated by pelicans on bicycles

simonwillison.net

21–30 of 244 posts

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#21

My biggest gripe is that he's comparing probabilistic models (LLMs) by a single sample. You wouldn't compare different random number generators by taking one sample from each and then concluding that generator 5 generates the highest numbers... Would be nicer to run the comparison with 10 images (or more) for each LLM and then average.

You are right, but the companies making these models invest a lot of effort in marketing them as anything but probabilistic, i.e. making people think that these models work discretely like humans. In that case we'd expect a human with perfect drawing skills and perfect knowledge about bikes and birds to output such a simple drawing correctly 100% of the time. In any case, even if a model is probabilistic, if it had c…

> In that case we'd expect a human with perfect drawing skills and perfect knowledge about bikes and birds to output such a simple drawing correctly 100% of the time.

Look upon these works, ye mighty, and despair: https://www.gianlucagimini.it/portfolio-item/velocipedia/

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#22
post #11
post #9

Earlier quoted context omitted.

That seems to be a completely inappropriate use case? I would not hire a blind artist or a deaf musician.

It's a proxy for abstract designing, like writing software or designing in a parametric CAD. Most the non-math design work of applied engineering AFAIK falls under the umbrella that's tested with the pelican riding the bicycle. You have to make a mental model and then turn it into applicable instructions. Program code/SVG markup/parametric CAD instructions don't really differ in that aspect.

I would not assume that this methodology applies to applied engineering, as a former actual real tangible meat space engineer. Things are a little nuanced and the nuances come from a combination of communication and experience, neither of which any LLM has any insight into at all. It's not out there on the internet to train it with and it's not even easy to put it into abstract terms which can be used as training data. And engineering itself in isolation doesn't exist - there is a whole world around it.

Ergo no you can't just say throw a bicycle into an LLM and a parametric model drops out into solidworks, then a machine makes it. And everyone buys it. That is the hope really isn't it? You end up with a useless shitty bike with a shit pelican on it.

The biggest problem we have in the LLM space is the fact that no one really knows any of the proposed use cases enough and neither does anyone being told that it works for the use cases.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#23
Here’s the spot where we see who’s TL;DR…

> Claude 4 will rat you out to the feds!

>If you expose it to evidence of malfeasance in your company, and you tell it it should act ethically, and you give it the ability to send email, it’ll rat you out.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#24
post #22
post #11

Earlier quoted context omitted.

It's a proxy for abstract designing, like writing software or designing in a parametric CAD. Most the non-math design work of applied engineering AFAIK falls under the umbrella that's tested with the pelican riding the bicycle. You have to make a mental model and then turn it into applicable instructions. Program code/SVG markup/parametric CAD instructions don't really differ in that aspect.

I would not assume that this methodology applies to applied engineering, as a former actual real tangible meat space engineer. Things are a little nuanced and the nuances come from a combination of communication and experience, neither of which any LLM has any insight into at all. It's not out there on the internet to train it with and it's not even easy to put it into abstract terms which can be used as training dat…

https://www.solidworks.com/lp/evolve-your-design-workflows-a...

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#25

Enjoyable write-up, but why is Qwen 3 conspicuously absent? It was a really strong release, especially the fine-grained MoE which is unlike anything that’s come before (in terms of capability and speed on consumer hardware).

Cut for time - qwen3 was pelican tested too https://simonwillison.net/2025/Apr/29/qwen-3/

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#26
post #8

Earlier quoted context omitted.

it depends on quality you need and your budget

Ah yes the race to the bottom argument.

When I was at university, they got some people from industry to talk to us all about our CVs and how to do interviews.

My CV had a stupid cliché, "committed to quality", which they correctly picked up on — "What do you mean?" one of them asked me, directly.

I thought this meant I was focussed on being the best. He didn't like this answer.

His example, blurred by 20 years of my imperfect human memory, was to ask me which is better: a Porsche, or a go-kart. Now, obviously (or I wouldn't be saying this), Porsche was a trick answer. Less obviously is that both were trick answers, because their point was that the question was under-specified — quality is the match between the product and what the user actually wants, so if the user is a 10 year old who physically isn't big enough to sit in a real car's driver's seat and just wants to rush down a hill or along a track, none of "quality" stuff that makes a Porsche a Porsche is of any relevance at all, but what does matter is the stuff that makes a go-kart into a go-kart… one of which is the affordability.

LLMs are go-karts of the mind. Sometimes that's all you need.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#27

Here’s the spot where we see who’s TL;DR… > Claude 4 will rat you out to the feds! >If you expose it to evidence of malfeasance in your company, and you tell it it should act ethically, and you give it the ability to send email, it’ll rat you out.

I'd say that's too short.

> But it’s not just Claude. Theo Browne put together a new benchmark called SnitchBench, inspired by the Claude 4 System Card.

> It turns out nearly all of the models do the same thing.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#28
post #22
post #11

Earlier quoted context omitted.

It's a proxy for abstract designing, like writing software or designing in a parametric CAD. Most the non-math design work of applied engineering AFAIK falls under the umbrella that's tested with the pelican riding the bicycle. You have to make a mental model and then turn it into applicable instructions. Program code/SVG markup/parametric CAD instructions don't really differ in that aspect.

I would not assume that this methodology applies to applied engineering, as a former actual real tangible meat space engineer. Things are a little nuanced and the nuances come from a combination of communication and experience, neither of which any LLM has any insight into at all. It's not out there on the internet to train it with and it's not even easy to put it into abstract terms which can be used as training dat…

I don't think any of that matters, CEOs will decide to use it anyway.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#29

My biggest gripe is that he's comparing probabilistic models (LLMs) by a single sample. You wouldn't compare different random number generators by taking one sample from each and then concluding that generator 5 generates the highest numbers... Would be nicer to run the comparison with 10 images (or more) for each LLM and then average.

It might not be 100% clear from the writing but this benchmark is mainly intended as a joke - I built a talk around it because it's a great way to make the last six months of model releases a lot more entertaining.

I've been considering an expanded version of this where each model outputs ten images, then a vision model helps pick the "best" of those to represent that model in a further competition with other models.

(Then I would also expand the judging panel to three vision LLMs from different model families which vote on each round... partly because it will be interesting to track cases where the judges disagree.)

I'm not sure if it's worth me doing that though since the whole "benchmark" is pretty silly. I'm on the fence.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#30
post #9

Earlier quoted context omitted.

That seems to be a completely inappropriate use case? I would not hire a blind artist or a deaf musician.

The point is about exploring the capabilities of the model. Like asking you to draw a 2D projection of 4D sphere intersected with a 4D torus or something.

[deleted]
Post reply on HN