Live data from Hacker News

The last six months in LLMs, illustrated by pelicans on bicycles

simonwillison.net

231–240 of 244 posts

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#231

Earlier quoted context omitted.

Not sure if this is sarcasm or sincere, but I will take it as sincere haha. I came back to work from parental leave and everyone had that same Studio Ghiblized image as their Slack photo, and I had no idea why. It turns out you really can unplug from social media and not miss anything of value: if it’s a big enough deal you will find out from another channel.

Why does everyone keep calling news "social media"? Have I missed a trend? Knowing what my friend Steve is up to is social media, knowing what AI is up to is news.

I'm afraid a lot of Americans consume the news like they consume sports media. They root for their team and select a news stream that presents them with the most favorable coverage.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#232

Earlier quoted context omitted.

Why does everyone keep calling news "social media"? Have I missed a trend? Knowing what my friend Steve is up to is social media, knowing what AI is up to is news.

I'm afraid a lot of Americans consume the news like they consume sports media. They root for their team and select a news stream that presents them with the most favorable coverage.

As a non-American, I can assure you that's pretty much everywhere.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#233
I think its hilarious how humans can make mistakes interpreting the crazy drawings : He says "I like how it solved the problem of pelicans not fitting on bicycles by adding a second smaller bicycle to the stack."

no... that is an attempt at it actually drawing the pedals, and putting the pelicans feet right on the pedals!

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#234
post #200
post #197

Earlier quoted context omitted.

Is that information about their low conversion rates from credible sources?

It's quite hard to say for sure, and I will prefix my comment by saying his blog posts are very long and quite doomerist about LLMs, but he makes a decent case about OpenAI financials: https://www.wheresyoured.at/wheres-the-money/ https://www.wheresyoured.at/openai-is-a-systemic-risk-to-the... A very solid argument is like that against propaganda: it's not so much about what is being said but what about isn't. OpenAI…

All very fair caveats/heads up about Ed Zitron, but just for context for others: he is an actual journalist that has been in the tech space for a long time, and has been critical of lots of large figures in tech for a long time. He has a cohesive thesis around the tech industry, so his thoughts on AI/LLMs aren't out of nowhere and disconnected.

Basically, it's one of those things you may read and find that, all things considered, you don't agree with the conclusions, but there's real substance there and you'll probably benefit from reading a few of his articles.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#235

Earlier quoted context omitted.

This is why things like the ARC Prize are better ways of approaching this: https://arcprize.org

Well, ARC-1 did not end well for the competitors of tech giants and it’s very unclear that ARC-2 won’t follow the same trajectory.

This doesn’t make ARC a bad benchmark. Tech giants will have a significant advantage in any benchmark they are interested in, _especially_ if the benchmark correlates with true general intelligence.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#236
post #202

So the only bird slightly resembling a pelican beak was drawn by gemini 2.5 pro. In general, none of the output resembles a pelican enough so you could separate it from "a bird". OP seem to ignore that pelican has a distinct look when evaluating these doodles.

The pelican's distinct look - and the fact that none of the models can capture it - is the whole point.

> The pelican's distinct look - and the fact that none of the models can capture it - is the whole point.

You didn't even mention the beak or the lack of similarities in your blog.

Your text is centered around this rather peculiar statement:

> Most importantly: pelicans can’t ride bicycles.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#237
post #202

Earlier quoted context omitted.

The pelican's distinct look - and the fact that none of the models can capture it - is the whole point.

> The pelican's distinct look - and the fact that none of the models can capture it - is the whole point. You didn't even mention the beak or the lack of similarities in your blog. Your text is centered around this rather peculiar statement: > Most importantly: pelicans can’t ride bicycles.

The blog was my attempt to capture the key ideas from the talk, which was full of jokes that don't come across as well in text as they do out loud.

"Pelicans can't ride bicycles" is a good joke.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#238

Earlier quoted context omitted.

You're right, nothing has value unless someone figures out how to make money with it. Except OpenAI, apparently, because the fact that people buy ChatGPT to make images doesn't seem to count as a commercial use case.

OpenAI is not profitable and we don't know if it ever will be.

OpenAI is not profitable because it is spending resources into moving forward and training new models and creating new tools.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#239
post #214
post #185

Earlier quoted context omitted.

I missed it until this thread. I think I’m proud of myself.

You're one of today's lucky 10.000 https://xkcd.com/1053/

I’m not, I still don’t know what it is, just that it was some kind of fad.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#240
post #29

My biggest gripe is that he's comparing probabilistic models (LLMs) by a single sample. You wouldn't compare different random number generators by taking one sample from each and then concluding that generator 5 generates the highest numbers... Would be nicer to run the comparison with 10 images (or more) for each LLM and then average.

It might not be 100% clear from the writing but this benchmark is mainly intended as a joke - I built a talk around it because it's a great way to make the last six months of model releases a lot more entertaining. I've been considering an expanded version of this where each model outputs ten images, then a vision model helps pick the "best" of those to represent that model in a further competition with other models.…

I'd be really interested in evaluating the evaluations of different models. At work, I maintain our internal LLM benchmarks for content generation. We've always used human raters from MTurk, and the Elo rankings generally match what you'd expect. I'm looking at our options for having LLMs do the evaluating.

In your case, it would be neat to have a bunch of different models (and maybe MTurk) pick the winners of each head-to-head matchup and then compare how stable the Elo scores are between evaluators.

Post reply on HN