Live data from Hacker News

The last six months in LLMs, illustrated by pelicans on bicycles

simonwillison.net

201–210 of 244 posts

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#201
So the only bird slightly resembling a pelican beak was drawn by gemini 2.5 pro. In general, none of the output resembles a pelican enough so you could separate it from "a bird".

OP seem to ignore that pelican has a distinct look when evaluating these doodles.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#202

So the only bird slightly resembling a pelican beak was drawn by gemini 2.5 pro. In general, none of the output resembles a pelican enough so you could separate it from "a bird". OP seem to ignore that pelican has a distinct look when evaluating these doodles.

The pelican's distinct look - and the fact that none of the models can capture it - is the whole point.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#203
post #96

Earlier quoted context omitted.

The big trend was around the ghiblification of images. Those images were everywhere for a period of time.

Yeah, but so were the bored ape NFTs - none of these ephemeral fads are any indication of quality, longevity, legitimacy, or interest.

I just don't understand how people can see "100 million signups in a week" and immediately dismiss it. We're not talking about fidget spinners. I don't get why this sentiment is so common here on HackerNews. It's become a running joke in other online spaces, "HackerNews commenters keep saying that AI is a nothingburger." It's just a groupthink thing I guess, a kneejerk response.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#204

> I’ve been feeling pretty good about my benchmark! It should stay useful for a long time... provided none of the big AI labs catch on. > And then I saw this in the Google I/O keynote a few weeks ago, in a blink and you’ll miss it moment! There’s a pelican riding a bicycle! They’re on to me. I’m going to have to switch to something else. Yeah this touches on an issue that makes it very difficult to have a discussion…

This is why things like the ARC Prize are better ways of approaching this: https://arcprize.org

Well, ARC-1 did not end well for the competitors of tech giants and it’s very unclear that ARC-2 won’t follow the same trajectory.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#205
post #96

Earlier quoted context omitted.

Yeah, but so were the bored ape NFTs - none of these ephemeral fads are any indication of quality, longevity, legitimacy, or interest.

I just don't understand how people can see "100 million signups in a week" and immediately dismiss it. We're not talking about fidget spinners. I don't get why this sentiment is so common here on HackerNews. It's become a running joke in other online spaces, "HackerNews commenters keep saying that AI is a nothingburger." It's just a groupthink thing I guess, a kneejerk response.

I assume, when people dismiss it, they are not looking at it through the business lens and the 100m user signups KPI, but they are dismissing it on technical grounds, as an LLM is just a very big statistical database which seems incapable of solving problems beyond (impressive looking) text/image/video generation.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#206

> This was one of the most successful product launches of all time. They signed up 100 million new user accounts in a week! They had a single hour where they signed up a million new accounts, as this thing kept on going viral again and again and again. Awkwardly, I never heard of it until now. I was aware that at some point they added ability to generate images to the app, but I never realized it was a major thing (p…

To be clear: they already had image generation in ChatGPT, but this was a MUCH better one than what they had previously. Even for you with your stable diffusion app, it would be a significant upgrade. Not just because of image quality, but because it can actually generate coherent images and follow instructions.

As impressive as it is, for some uses it still is worse than a local SD model. It will refuse to generate named anime characters (because of copyright, or because it just doesn't know them, even not particularly obscure ones) for example. Or obviously anything even remotely spicy. As someone who mostly uses image generation to amuse myself (and not to post it, where copyright might matter) it's honestly somewhat disappointing. But I don't expect any of the major AI companies to release anything without excessive guardrails.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#207
post #205

Earlier quoted context omitted.

I just don't understand how people can see "100 million signups in a week" and immediately dismiss it. We're not talking about fidget spinners. I don't get why this sentiment is so common here on HackerNews. It's become a running joke in other online spaces, "HackerNews commenters keep saying that AI is a nothingburger." It's just a groupthink thing I guess, a kneejerk response.

I assume, when people dismiss it, they are not looking at it through the business lens and the 100m user signups KPI, but they are dismissing it on technical grounds, as an LLM is just a very big statistical database which seems incapable of solving problems beyond (impressive looking) text/image/video generation.

Makes sense. Although I think that's an error. TikTok is "just" a video sharing site. Joe Rogan is "just" a podcaster. Dumb things that affect lots of people are important.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#208
post #115
post #96

Earlier quoted context omitted.

Yeah, but so were the bored ape NFTs - none of these ephemeral fads are any indication of quality, longevity, legitimacy, or interest.

It’s hard to think of a worse analogy TBH. My wife is using ChatGPT to change photos (still is to this day), she didn’t use it or any other LLM until that feature hit. It is a fad, but it’s also a very useful tool. Ape NFTs are… ape NFTs. Useless. Pointless. Negative value for most people.

I would note that I was replying to a comment about the 'big trend of ghiblification' of images.

Reproducing a certain style of image has been a regular fad since profile pictures became a thing sometime last century.

I was not meaning to suggest that large language & diffusion models are fads.

(I do think their capabilities are poorly understood and/or over-estimated by non-technical and some technical people alike, but that invites a more nuanced discussion.)

While I'm sure your wife is getting good value out of the system, whether it's a better fit for purpose, produces a better quality, or provides a more satisfying workflow -- than say a decent free photo editor -- or whether other tools were tried but determined to be too limited or difficult, etc -- only you or her could say. It does feel like a small sample set, though.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#209
post #123

Earlier quoted context omitted.

> Ask 100 random people to draw a bike and in 10 minutes and they’ll on average suck while still beating the LLM’s here. Y'see, this is a prime example of what I meant with ""Average human" is a much lower bar than most people want to believe, mainly because most of us are average on most skills, and also overestimate our own competence". An expert artist can spend 10 minutes and end up with a brief sketch of a bike.…

> A normal person spending as much time as they like gets you the pictures that I linked to in the previous post, because they don't really know what a bike is. 45 examples of what normal people think a bike looks like: https://www.gianlucagimini.it/portfolio-item/velocipedia/ A normal person given the ability to consult a picture of a bike while drawing will do much better. An LLM agent can effectively refresh its m…

> A normal person given the ability to consult a picture of a bike while drawing will do much better. An LLM agent can effectively refresh its memory (or attempt to look up information on the Internet) any time it wants.

Some models can when allowed to, but I don't belive Simon Willson was testing that?

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#210
post #104

Earlier quoted context omitted.

If we try really hard, I think we can make an exhaustive list of what viral fads on the internet are not. You made a small start. none of these ephemeral fads are any indication of quality, longevity, legitimacy, interest, substance, endurance, prestige, relevance, credibility, allure, staying-power, refinement, or depth.

100 million people didn’t sign up to make that one image meme and then never use it again. That many signups is impressive no matter what. The attempts to downplay every aspect of LLM popularity are getting really tiresome.

> 100 million people didn’t sign up to make that one image meme and then never use it again.

Source? They did exactly that.

Post reply on HN