Live data from Hacker News

The last six months in LLMs, illustrated by pelicans on bicycles

simonwillison.net

151–160 of 244 posts

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#151
post #115

Earlier quoted context omitted.

It’s hard to think of a worse analogy TBH. My wife is using ChatGPT to change photos (still is to this day), she didn’t use it or any other LLM until that feature hit. It is a fad, but it’s also a very useful tool. Ape NFTs are… ape NFTs. Useless. Pointless. Negative value for most people.

"My wife is using ChatGPT to change photos (still is to this day), she didn’t use it or any other LLM until that feature hit." This is deja vu, except instead of ChatGPT to edit photos it was instagram a decade ago.

You either haven’t tried it or are just trolling.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#152

> This was one of the most successful product launches of all time. They signed up 100 million new user accounts in a week! They had a single hour where they signed up a million new accounts, as this thing kept on going viral again and again and again. Awkwardly, I never heard of it until now. I was aware that at some point they added ability to generate images to the app, but I never realized it was a major thing (p…

Have you missed how everyone was Ghiblifying everything?

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#153
post #92
post #79

Earlier quoted context omitted.

until they start targeting this benchmark

Right, that was the closing joke for the talk.

It is funny to think that a hundred years in the future there may be some vestigial area of the models’ networks that’s still tuned to drawing pelicans on bicycles.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#154

My biggest gripe is that he's comparing probabilistic models (LLMs) by a single sample. You wouldn't compare different random number generators by taking one sample from each and then concluding that generator 5 generates the highest numbers... Would be nicer to run the comparison with 10 images (or more) for each LLM and then average.

You are right, but the companies making these models invest a lot of effort in marketing them as anything but probabilistic, i.e. making people think that these models work discretely like humans. In that case we'd expect a human with perfect drawing skills and perfect knowledge about bikes and birds to output such a simple drawing correctly 100% of the time. In any case, even if a model is probabilistic, if it had c…

> work discretely like humans

What kind of humans are you surrounded by?

Ask any human to write 3 sentences about a specific topic. Then ask them the same exact question next day. They will not write the same 3 sentences.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#155
post #137
post #123

Earlier quoted context omitted.

> Ask 100 random people to draw a bike and in 10 minutes and they’ll on average suck while still beating the LLM’s here. Y'see, this is a prime example of what I meant with ""Average human" is a much lower bar than most people want to believe, mainly because most of us are average on most skills, and also overestimate our own competence". An expert artist can spend 10 minutes and end up with a brief sketch of a bike.…

As an objective criteria what percentage include peddles and a chain connecting one of the wheels? I quickly found a dozen and stopped counting. Now do the same for those LLM images and it’s clear humans win. > ""Average human" is a much lower bar than most people want to believe I have some basis for comparison. I’ve seen 6 years olds draw better bikes than those LLM’s. Look through that list again the worst example…

The 88 cent one was the most expensive almost my an order of magnitude. Most of these cost less than a cent to generate - that's why I highlighted the price on the o1 pro output.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#156
post #131

Earlier quoted context omitted.

I had the LLM write a bash script for me that used my https://shot-scraper.datasette.io/ tool - on the basis that it was a neat opportunity to demonstrate another of my own projects. And honestly, even with LLM assistance getting Image Magick to output a 1200x600 image with two SVGs next to each other that are correctly resized to fill their half of the image sounds pretty tricky. Probably easier (for Claude) to achi…

Isn't "left or right" _followed_ by rationale asking it to rationalize it's 1 word answer - I thought we need to get AI to do the chain of though _before_ giving it's answer for it to be more accurate?

Yes it is - I would likely have gotten better results if I'd asked for the rationale first.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#157
post #155
post #137

Earlier quoted context omitted.

As an objective criteria what percentage include peddles and a chain connecting one of the wheels? I quickly found a dozen and stopped counting. Now do the same for those LLM images and it’s clear humans win. > ""Average human" is a much lower bar than most people want to believe I have some basis for comparison. I’ve seen 6 years olds draw better bikes than those LLM’s. Look through that list again the worst example…

The 88 cent one was the most expensive almost my an order of magnitude. Most of these cost less than a cent to generate - that's why I highlighted the price on the o1 pro output.

Yes, but if you’re averaging cheap and expensive options the expensive ones make a significant difference. Cheaper is bound by 0 so it can’t differ as much from the average.

Also, when you’re talking about how cheap something is, including the price makes sense. I had no idea on many of those models.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#158

> This was one of the most successful product launches of all time. They signed up 100 million new user accounts in a week! They had a single hour where they signed up a million new accounts, as this thing kept on going viral again and again and again. Awkwardly, I never heard of it until now. I was aware that at some point they added ability to generate images to the app, but I never realized it was a major thing (p…

Congratulations, you are almost fully unplugged from social media. This product launch was a huge mainstream event; for a few days GPT generated images completely dominated mainstream social media.

Not sure if this is sarcasm or sincere, but I will take it as sincere haha. I came back to work from parental leave and everyone had that same Studio Ghiblized image as their Slack photo, and I had no idea why. It turns out you really can unplug from social media and not miss anything of value: if it’s a big enough deal you will find out from another channel.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#159
post #157
post #155

Earlier quoted context omitted.

The 88 cent one was the most expensive almost my an order of magnitude. Most of these cost less than a cent to generate - that's why I highlighted the price on the o1 pro output.

Yes, but if you’re averaging cheap and expensive options the expensive ones make a significant difference. Cheaper is bound by 0 so it can’t differ as much from the average. Also, when you’re talking about how cheap something is, including the price makes sense. I had no idea on many of those models.

If you're interested, you can get cost estimates from my pricing calculator site here: https://www.llm-prices.com/#it=11&ot=1200

That link seeds it with 11 input tokens and 1200 output tokens - 11 input tokens is what most models use for "Generate an SVG of a pelican riding a bicycle" and 1200 is the number of output tokens used for some of the larger outputs.

Click on different models to see estimated prices. They range from 0.0168 cents for Amazon Nova Micro (that's less than 2/100ths of a cent) up to 72 cents for o1-pro.

The most expensive model most people would consider is Claude 4 Opus, at 9 cents.

GPT-4o is the upper end of the most common prices, at 1.2 cents.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#160
post #159
post #157

Earlier quoted context omitted.

Yes, but if you’re averaging cheap and expensive options the expensive ones make a significant difference. Cheaper is bound by 0 so it can’t differ as much from the average. Also, when you’re talking about how cheap something is, including the price makes sense. I had no idea on many of those models.

If you're interested, you can get cost estimates from my pricing calculator site here: https://www.llm-prices.com/#it=11&ot=1200 That link seeds it with 11 input tokens and 1200 output tokens - 11 input tokens is what most models use for "Generate an SVG of a pelican riding a bicycle" and 1200 is the number of output tokens used for some of the larger outputs. Click on different models to see estimated prices. They r…

Thanks
Post reply on HN