Live data from Hacker News

The last six months in LLMs, illustrated by pelicans on bicycles

simonwillison.net

241–244 of 244 posts

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#241

Earlier quoted context omitted.

I'd say definitely do not do that. That would make the benchmark look more serious while still being problematic for knowledge cutoff reasons. Your prompt has become popular even outside your blog, so the odds of some SVG pelicans on bicycles making it into the training data have been going up and up. Karpathy used it as an example in a recent interview: https://www.msn.com/en-in/health/other/ai-expert-asks-grok-3...

Yeah, Simon needs to release a new benchmark under a pen name, like Stephen King did with Richard Bachman.

Richard Bachman, you say? https://chatgpt.com/share/684c3f20-575c-800a-9ea2-889dd3deaf...

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#242
post #50
post #29

Earlier quoted context omitted.

It might not be 100% clear from the writing but this benchmark is mainly intended as a joke - I built a talk around it because it's a great way to make the last six months of model releases a lot more entertaining. I've been considering an expanded version of this where each model outputs ten images, then a vision model helps pick the "best" of those to represent that model in a further competition with other models.…

Even if it is a joke, having a consistent methodology is useful. I did it for about a year with my own private benchmark of reasoning type questions that I always applied to each new open model that came out. Run it once and you get a random sample of performance. Got unlucky, or got lucky? So what. That's the experimental protocol. Running things a bunch of times and cherry picking the best ones adds human bias, and…

Another advantage is you can easily include deprecated models in your comparisons. I maintain our internal LLM rankings at work. Since the prompts have remained the same, I can do things like compare the latest Gemini Pro to the original Bard.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#243
post #171

> I’ve been feeling pretty good about my benchmark! It should stay useful for a long time... provided none of the big AI labs catch on. > And then I saw this in the Google I/O keynote a few weeks ago, in a blink and you’ll miss it moment! There’s a pelican riding a bicycle! They’re on to me. I’m going to have to switch to something else. Yeah this touches on an issue that makes it very difficult to have a discussion…

Honestly, if my stupid pelican riding a bicycle benchmark becomes influential enough that AI labs waste their time optimizing for it and produce really beautiful pelican illustrations I will consider that a huge personal win.

"personal" doing a lot of work there :-)

(And I'd be envious of your impact, of course)

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#244

I really enjoy Simon’s work in this space. I’ve read almost every blog post they’ve posted on this and I love seeing them poke and prod the models to see what pops out. The CLI tools are all very easy to use and complement each other nicely all without trying to do too much by themselves. And at the end of the day, it’s just so much fun to see someone else having so much fun. He’s like a kid in a candy store and that…

I also like what he writes and the way he does it.
Post reply on HN