Live data from Hacker News

The last six months in LLMs, illustrated by pelicans on bicycles

simonwillison.net

51–60 of 244 posts

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#51
post #50
post #29

Earlier quoted context omitted.

It might not be 100% clear from the writing but this benchmark is mainly intended as a joke - I built a talk around it because it's a great way to make the last six months of model releases a lot more entertaining. I've been considering an expanded version of this where each model outputs ten images, then a vision model helps pick the "best" of those to represent that model in a further competition with other models.…

Even if it is a joke, having a consistent methodology is useful. I did it for about a year with my own private benchmark of reasoning type questions that I always applied to each new open model that came out. Run it once and you get a random sample of performance. Got unlucky, or got lucky? So what. That's the experimental protocol. Running things a bunch of times and cherry picking the best ones adds human bias, and…

It wasn't until I put these slides together that I realized quite how well my joke benchmark correlates with actual model performance - the "better" models genuinely do appear to draw better pelicans and I don't really understand why!

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#54
post #33

Earlier quoted context omitted.

The bicycle are still very far from actual ones.

I think the most recent Gemini Pro bicycle may be the best yet - the red frame is genuinely the right shape.

The pelican, on the other hand...

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#55
post #43
post #26

Earlier quoted context omitted.

When I was at university, they got some people from industry to talk to us all about our CVs and how to do interviews. My CV had a stupid cliché, "committed to quality", which they correctly picked up on — "What do you mean?" one of them asked me, directly. I thought this meant I was focussed on being the best. He didn't like this answer. His example, blurred by 20 years of my imperfect human memory, was to ask me wh…

I disagree. Quality depends on your market position and what you are bringing to the market. Thus I would start with market conditions and work back to quality. If you can't reach your standards in the market then you shouldn't enter it. And if your standards are poor, you should be ashamed. Go kart or porsche is irrelevant.

> Quality depends on your market position and what you are bringing to the market.

That's the point.

The market for go-karts does not support Porche.

If you bring a Porche sales team to a go-kart race, nobody will be interested.

Porche doesn't care about this market. It goes both ways: this market doesn't care about Porche, either.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#57

My biggest gripe is that he's comparing probabilistic models (LLMs) by a single sample. You wouldn't compare different random number generators by taking one sample from each and then concluding that generator 5 generates the highest numbers... Would be nicer to run the comparison with 10 images (or more) for each LLM and then average.

I think you mean non-deterministic, instead of probabilistic.

And there is no reason that these models need to be non-deterministic.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#58
post #51
post #50

Earlier quoted context omitted.

Even if it is a joke, having a consistent methodology is useful. I did it for about a year with my own private benchmark of reasoning type questions that I always applied to each new open model that came out. Run it once and you get a random sample of performance. Got unlucky, or got lucky? So what. That's the experimental protocol. Running things a bunch of times and cherry picking the best ones adds human bias, and…

It wasn't until I put these slides together that I realized quite how well my joke benchmark correlates with actual model performance - the "better" models genuinely do appear to draw better pelicans and I don't really understand why!

How did the pelicans of point releases of V3 and of R1 (R1-0528) do compared to the original versions of the models?

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#60

Earlier quoted context omitted.

You are right, but the companies making these models invest a lot of effort in marketing them as anything but probabilistic, i.e. making people think that these models work discretely like humans. In that case we'd expect a human with perfect drawing skills and perfect knowledge about bikes and birds to output such a simple drawing correctly 100% of the time. In any case, even if a model is probabilistic, if it had c…

Humans absolutely do not work discretely.

They probably meant deterministically as opposed to probabilistically. Which also humans dont work like that :)
Post reply on HN