Live data from Hacker News

The last six months in LLMs, illustrated by pelicans on bicycles

simonwillison.net

101–110 of 244 posts

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#101
post #21

Earlier quoted context omitted.

> In that case we'd expect a human with perfect drawing skills and perfect knowledge about bikes and birds to output such a simple drawing correctly 100% of the time. Look upon these works, ye mighty, and despair: https://www.gianlucagimini.it/portfolio-item/velocipedia/

You claim those are drawn by people with "perfect knowledge about bikes" and "perfect drawing skills"?

More that "these models work … like humans" (discretely or otherwise) does not imply the quotation.

Most humans do not have perfect drawing skills and perfect knowledge about bikes and birds, they do not output such a simple drawing correctly 100% of the time.

"Average human" is a much lower bar than most people want to believe, mainly because most of us are average on most skills, and also overestimate our own competence — the modal human has just a handful of things they're good at, and one of those is the language they use, another is their day job.

Most of us can't draw, and demonstrably can't remember (or figure out from first principles) how a bike works. But this also applies to "smart" subsets of the population: physicists have https://xkcd.com/793/, and there's this famous rocket scientist who weighed in on rescuing kids from a flooded cave, they come up with some nonsense about a submarine.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#102
post #29

My biggest gripe is that he's comparing probabilistic models (LLMs) by a single sample. You wouldn't compare different random number generators by taking one sample from each and then concluding that generator 5 generates the highest numbers... Would be nicer to run the comparison with 10 images (or more) for each LLM and then average.

It might not be 100% clear from the writing but this benchmark is mainly intended as a joke - I built a talk around it because it's a great way to make the last six months of model releases a lot more entertaining. I've been considering an expanded version of this where each model outputs ten images, then a vision model helps pick the "best" of those to represent that model in a further competition with other models.…

Joke or not, it still correlates much better with my own subjective experiences of the models than LM Arena!

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#104
post #96

Earlier quoted context omitted.

The big trend was around the ghiblification of images. Those images were everywhere for a period of time.

Yeah, but so were the bored ape NFTs - none of these ephemeral fads are any indication of quality, longevity, legitimacy, or interest.

If we try really hard, I think we can make an exhaustive list of what viral fads on the internet are not. You made a small start.

none of these ephemeral fads are any indication of quality, longevity, legitimacy, interest, substance, endurance, prestige, relevance, credibility, allure, staying-power, refinement, or depth.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#105
post #51
post #50

Earlier quoted context omitted.

Even if it is a joke, having a consistent methodology is useful. I did it for about a year with my own private benchmark of reasoning type questions that I always applied to each new open model that came out. Run it once and you get a random sample of performance. Got unlucky, or got lucky? So what. That's the experimental protocol. Running things a bunch of times and cherry picking the best ones adds human bias, and…

It wasn't until I put these slides together that I realized quite how well my joke benchmark correlates with actual model performance - the "better" models genuinely do appear to draw better pelicans and I don't really understand why!

Well, the most likely single random sample would be a “representative” one :)

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#106

Earlier quoted context omitted.

Except this went very mainstream. Lots of turn myself into a muppet, what is the human equivalent for my dog, etc. TikTok is all over this. It really is incredible.

The big trend was around the ghiblification of images. Those images were everywhere for a period of time.

They still are. Instagram is full of accounts posting gpt-generated cartoons (and now veo3 videos). I’ve been tracking the image generation space from day one, and it never stuck like this before

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#107
post #106

Earlier quoted context omitted.

The big trend was around the ghiblification of images. Those images were everywhere for a period of time.

They still are. Instagram is full of accounts posting gpt-generated cartoons (and now veo3 videos). I’ve been tracking the image generation space from day one, and it never stuck like this before

Anecdotally, I've had several conversations with people way outside the hyper-online demographic who have been really enjoying the new ChatGPT image generation - using it for cartoon photos of their kids, to create custom birthday cards etc.

I think it's broken out into mainstream adoption and is going to stay there.

It reminds me a little of Napster. The Napster UI was terrible, but it let people do something they had never been able to do before: listen to any piece of music ever released, on-demand. As a result people with almost no interest in technology at all were learning how to use it.

Most people have never had the ability to turn a photo of their kids into a cute cartoon before, and it turns out that's something they really want to be able to do.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#109
post #107
post #106

Earlier quoted context omitted.

They still are. Instagram is full of accounts posting gpt-generated cartoons (and now veo3 videos). I’ve been tracking the image generation space from day one, and it never stuck like this before

Anecdotally, I've had several conversations with people way outside the hyper-online demographic who have been really enjoying the new ChatGPT image generation - using it for cartoon photos of their kids, to create custom birthday cards etc. I think it's broken out into mainstream adoption and is going to stay there. It reminds me a little of Napster. The Napster UI was terrible, but it let people do something they h…

Definitely. It’s not just online either - half the billboards I see now are AI. The posters at school. The “we’re hiring!” ad at the local McDonalds. It’s 100x cheaper and faster than any alternative (stock images, hiring an editor or illustrator, etc), and most non technical people can get exactly what they want in a single shot, these days.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#110
post #74

Earlier quoted context omitted.

Yeah, this is the problem with benchmarks where the questions/problems are public. They're valuable for some months, until it bleeds into the training set. I'm certain a lot of the "improvements" we're seeing are just benchmarks leaking into the training set.

That’s ok, once bicycle “riding” pelicans become normative, we can ask it for images of pelicans humping bicycles. The number of subject-verb-objects are near infinite. All are imaginable, but most are not plausible. A plausibility machine (LLM) will struggle with the implausible, until it can abstract well.

> The number of subject-verb-objects are near infinite. All are imaginable, but most are not plausible

Until there is enough unique/new subject-verb-objects examples/benchmarks so the trained model actually generalized it just like you did. (Public) Benchmarks needs to constantly evolve, otherwise they stop being useful.

Post reply on HN