Live data from Hacker News

The last six months in LLMs, illustrated by pelicans on bicycles

simonwillison.net

61–70 of 244 posts

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#61
post #27

Here’s the spot where we see who’s TL;DR… > Claude 4 will rat you out to the feds! >If you expose it to evidence of malfeasance in your company, and you tell it it should act ethically, and you give it the ability to send email, it’ll rat you out.

I'd say that's too short. > But it’s not just Claude. Theo Browne put together a new benchmark called SnitchBench, inspired by the Claude 4 System Card. > It turns out nearly all of the models do the same thing.

I totally agree, but I needed you to post the other half because of TL;DR…

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#63

As a control, he should go on fiver and have a human generate a pelican riding a bicycle, just to see what the eventual goal is.

Someone did this. Look at this sibling comment by ben_w https://news.ycombinator.com/item?id=44216284 about an old similar project.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#64
post #17

My biggest gripe is that he's comparing probabilistic models (LLMs) by a single sample. You wouldn't compare different random number generators by taking one sample from each and then concluding that generator 5 generates the highest numbers... Would be nicer to run the comparison with 10 images (or more) for each LLM and then average.

And by a sample that has become increasingly known as a benchmark. Newer training data will contain more articles like this one, which naturally improves the capabilities of an LLM to estimate what’s considered a good „pelican on a bike“.

Would it though? There really aren't that many valid answers to that question online. When this is talked about, we get more broken samples than reasonable ones. I feel like any talk about this actually sabotages future training a bit.

I actually don't think I've seen a single correct svg drawing for that prompt.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#65

My biggest gripe is that he's comparing probabilistic models (LLMs) by a single sample. You wouldn't compare different random number generators by taking one sample from each and then concluding that generator 5 generates the highest numbers... Would be nicer to run the comparison with 10 images (or more) for each LLM and then average.

My biggest gripe is that he outsourced evaluation of the pelicans to another LLM.

I get it was way easier to do and that doing it took pennies and no time. But I would have loved it if he'd tried alternate methods of judging and seen what the results were.

Other ways:

* wisdom of the crowds (have people vote on it)

* wisdom of the experts (send the pelican images to a few dozen artists or ornithologists)

* wisdom of the LLMs (use more than one LLM)

Would have been neat to see what the human consensus was and if it differed from the LLM consensus

Anyway, great talk!

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#66
post #29

Earlier quoted context omitted.

It might not be 100% clear from the writing but this benchmark is mainly intended as a joke - I built a talk around it because it's a great way to make the last six months of model releases a lot more entertaining. I've been considering an expanded version of this where each model outputs ten images, then a vision model helps pick the "best" of those to represent that model in a further competition with other models.…

Very nice talk, acceptable by general public and by AI agent as well. Any concerns about open source “AI celebrity talks” like yours can be used in contexts that would allow LLM models to optimize their market share in ways that we can’t imagine yet? Your talk might influence the funding of AI startups. #butterflyEffect

I welcome a VC funded pelican … anything! Clippy 2.0 maybe?

Simon, hope you are comfortable in your new role of AI Celebrity.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#67
post #2

My only take home is they are all terrible and I should hire a professional.

Before that, you might ask ChatGPT to create a vector image of a pelican riding a bicycle and then running the output through a PNG to SVG converter...

Result: https://www.dropbox.com/scl/fi/8b03yu5v58w0o5he1zayh/pelican...

These are tough benchmarks to trial reasoning by having it _write_ an SVG file by hand and understanding how it's to be written to achieve this. Even a professional would struggle with that! It's _not_ a benchmark to give an AI the best tools to actually do this.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#68
post #51
post #50

Earlier quoted context omitted.

Even if it is a joke, having a consistent methodology is useful. I did it for about a year with my own private benchmark of reasoning type questions that I always applied to each new open model that came out. Run it once and you get a random sample of performance. Got unlucky, or got lucky? So what. That's the experimental protocol. Running things a bunch of times and cherry picking the best ones adds human bias, and…

It wasn't until I put these slides together that I realized quite how well my joke benchmark correlates with actual model performance - the "better" models genuinely do appear to draw better pelicans and I don't really understand why!

I imagine the straightforward reason is that the “better” models are in fact significantly smarter in some tangible way, somehow.
Post reply on HN