Live data from Hacker News

The last six months in LLMs, illustrated by pelicans on bicycles

simonwillison.net

41–50 of 244 posts

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#41
post #17

My biggest gripe is that he's comparing probabilistic models (LLMs) by a single sample. You wouldn't compare different random number generators by taking one sample from each and then concluding that generator 5 generates the highest numbers... Would be nicer to run the comparison with 10 images (or more) for each LLM and then average.

And by a sample that has become increasingly known as a benchmark. Newer training data will contain more articles like this one, which naturally improves the capabilities of an LLM to estimate what’s considered a good „pelican on a bike“.

And that’s why he says he’s going to have to find a new benchmark.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#42
post #36
post #9

Earlier quoted context omitted.

That seems to be a completely inappropriate use case? I would not hire a blind artist or a deaf musician.

Yeah, that's part of the point of this. Getting a state of the art text generating LLM to generate SVG illustrations is an inappropriate application of them. It's a fun way to deflate the hype. Sure, your new LLM may have cost XX million to train and beat all the others on the benchmarks, but when you ask it to draw a pelican on a bicycle it still outputs total junk.

tried starting from an image:

https://chatgpt.com/share/684582a0-03cc-8006-b5b5-de51e5cd89...

lol: https://gemini.google.com/share/4d1746a234a8

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#43
post #26
post #8

Earlier quoted context omitted.

Ah yes the race to the bottom argument.

When I was at university, they got some people from industry to talk to us all about our CVs and how to do interviews. My CV had a stupid cliché, "committed to quality", which they correctly picked up on — "What do you mean?" one of them asked me, directly. I thought this meant I was focussed on being the best. He didn't like this answer. His example, blurred by 20 years of my imperfect human memory, was to ask me wh…

I disagree. Quality depends on your market position and what you are bringing to the market. Thus I would start with market conditions and work back to quality. If you can't reach your standards in the market then you shouldn't enter it. And if your standards are poor, you should be ashamed.

Go kart or porsche is irrelevant.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#44
post #28
post #22

Earlier quoted context omitted.

I would not assume that this methodology applies to applied engineering, as a former actual real tangible meat space engineer. Things are a little nuanced and the nuances come from a combination of communication and experience, neither of which any LLM has any insight into at all. It's not out there on the internet to train it with and it's not even easy to put it into abstract terms which can be used as training dat…

I don't think any of that matters, CEOs will decide to use it anyway.

This is sad but true.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#45
post #9

Earlier quoted context omitted.

Most of them are text-only models. Like asking a person born blind to draw a pelican, based on what they heard it looks like.

That seems to be a completely inappropriate use case? I would not hire a blind artist or a deaf musician.

Even Beethoven?

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#46

Here’s the spot where we see who’s TL;DR… > Claude 4 will rat you out to the feds! >If you expose it to evidence of malfeasance in your company, and you tell it it should act ethically, and you give it the ability to send email, it’ll rat you out.

I was looking at that and wondering about swatting via LLMs by malicious users.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#47
post #29

My biggest gripe is that he's comparing probabilistic models (LLMs) by a single sample. You wouldn't compare different random number generators by taking one sample from each and then concluding that generator 5 generates the highest numbers... Would be nicer to run the comparison with 10 images (or more) for each LLM and then average.

It might not be 100% clear from the writing but this benchmark is mainly intended as a joke - I built a talk around it because it's a great way to make the last six months of model releases a lot more entertaining. I've been considering an expanded version of this where each model outputs ten images, then a vision model helps pick the "best" of those to represent that model in a further competition with other models.…

Very nice talk, acceptable by general public and by AI agent as well.

Any concerns about open source “AI celebrity talks” like yours can be used in contexts that would allow LLM models to optimize their market share in ways that we can’t imagine yet?

Your talk might influence the funding of AI startups.

#butterflyEffect

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#48
post #22

Earlier quoted context omitted.

I would not assume that this methodology applies to applied engineering, as a former actual real tangible meat space engineer. Things are a little nuanced and the nuances come from a combination of communication and experience, neither of which any LLM has any insight into at all. It's not out there on the internet to train it with and it's not even easy to put it into abstract terms which can be used as training dat…

https://www.solidworks.com/lp/evolve-your-design-workflows-a...

Yeah good luck with that. Seriously.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#49
post #29

My biggest gripe is that he's comparing probabilistic models (LLMs) by a single sample. You wouldn't compare different random number generators by taking one sample from each and then concluding that generator 5 generates the highest numbers... Would be nicer to run the comparison with 10 images (or more) for each LLM and then average.

It might not be 100% clear from the writing but this benchmark is mainly intended as a joke - I built a talk around it because it's a great way to make the last six months of model releases a lot more entertaining. I've been considering an expanded version of this where each model outputs ten images, then a vision model helps pick the "best" of those to represent that model in a further competition with other models.…

I'd say definitely do not do that. That would make the benchmark look more serious while still being problematic for knowledge cutoff reasons. Your prompt has become popular even outside your blog, so the odds of some SVG pelicans on bicycles making it into the training data have been going up and up.

Karpathy used it as an example in a recent interview: https://www.msn.com/en-in/health/other/ai-expert-asks-grok-3...

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#50
post #29

My biggest gripe is that he's comparing probabilistic models (LLMs) by a single sample. You wouldn't compare different random number generators by taking one sample from each and then concluding that generator 5 generates the highest numbers... Would be nicer to run the comparison with 10 images (or more) for each LLM and then average.

It might not be 100% clear from the writing but this benchmark is mainly intended as a joke - I built a talk around it because it's a great way to make the last six months of model releases a lot more entertaining. I've been considering an expanded version of this where each model outputs ten images, then a vision model helps pick the "best" of those to represent that model in a further competition with other models.…

Even if it is a joke, having a consistent methodology is useful. I did it for about a year with my own private benchmark of reasoning type questions that I always applied to each new open model that came out. Run it once and you get a random sample of performance. Got unlucky, or got lucky? So what. That's the experimental protocol. Running things a bunch of times and cherry picking the best ones adds human bias, and complicates the steps.
Post reply on HN