Live data from Hacker News

The last six months in LLMs, illustrated by pelicans on bicycles

simonwillison.net

131–140 of 244 posts

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#131

> If you lost interest in local models—like I did eight months ago—it’s worth paying attention to them again. They’ve got good now! > As a power user of these tools, I want to stay in complete control of what the inputs are. Features like ChatGPT memory are taking that control away from me. You reap what you sow.... > I already have a tool I built called shot-scraper, a CLI app that lets me take screenshots of web pa…

I had the LLM write a bash script for me that used my https://shot-scraper.datasette.io/ tool - on the basis that it was a neat opportunity to demonstrate another of my own projects.

And honestly, even with LLM assistance getting Image Magick to output a 1200x600 image with two SVGs next to each other that are correctly resized to fill their half of the image sounds pretty tricky. Probably easier (for Claude) to achieve with HTML and CSS.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#132
post #117
post #108

Am I the only one who can't but see these attempts much like attempts of a kid learning to draw?

Yes. Kids don't draw that good of a line at the start. Here is better example of start https://encrypted-tbn0.gstatic.com/images?q=tbn:ANd9GcTfTfAA...

Have you tried giving a kid a vector-drawing tool?

I did that to my daughter when she was not even 6 years old. The results were somehow similar: https://photos.app.goo.gl/XSLnTEUkmtW2n7cX8

(Now she's much better, but prefers raster tools, e.g. https://www.deviantart.com/sofiac9/art/Ivy-with-riding-gear-...)

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#133
post #123
post #112

Earlier quoted context omitted.

It’s not that humans have perfect drawing skills, it’s that humans can judge their performance and get better over time. Ask 100 random people to draw a bike and in 10 minutes and they’ll on average suck while still beating the LLM’s here. Give em an incentive and 10 months and the average person is going to be able to make at least one quite decent drawing of a bike. The cost and speed advantage of LLM’s is real as…

> Ask 100 random people to draw a bike and in 10 minutes and they’ll on average suck while still beating the LLM’s here. Y'see, this is a prime example of what I meant with ""Average human" is a much lower bar than most people want to believe, mainly because most of us are average on most skills, and also overestimate our own competence". An expert artist can spend 10 minutes and end up with a brief sketch of a bike.…

> A normal person spending as much time as they like gets you the pictures that I linked to in the previous post, because they don't really know what a bike is. 45 examples of what normal people think a bike looks like: https://www.gianlucagimini.it/portfolio-item/velocipedia/

A normal person given the ability to consult a picture of a bike while drawing will do much better. An LLM agent can effectively refresh its memory (or attempt to look up information on the Internet) any time it wants.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#134
post #65

My biggest gripe is that he's comparing probabilistic models (LLMs) by a single sample. You wouldn't compare different random number generators by taking one sample from each and then concluding that generator 5 generates the highest numbers... Would be nicer to run the comparison with 10 images (or more) for each LLM and then average.

My biggest gripe is that he outsourced evaluation of the pelicans to another LLM. I get it was way easier to do and that doing it took pennies and no time. But I would have loved it if he'd tried alternate methods of judging and seen what the results were. Other ways: * wisdom of the crowds (have people vote on it) * wisdom of the experts (send the pelican images to a few dozen artists or ornithologists) * wisdom of…

It would have been interesting to see if the LLM that Claude judged worst would have attempted to justify itself....

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#135
Does anyone have any thoughts on privacy/safety regarding what he said about GPT memory.

I had heard of prompt injection already. But, this seems different, completely out of humans control. Like even when you consider web search functionality, he is actually right, more and more, users are losing control over context.

Is this dangerous atm? Do you think it will become more dangerous in the future when we chuck even more data into context?

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#136
post #73

If you would give a human the SVG documentation and ask to write an SVG, I think the results would be quite similar.

Lets give it a try, if you're willing to be the experiment subject :) The prompt is "Generate an SVG of a pelican riding a bicycle" and you're supposed to write it by hand, so no graphical editor. The specification is here: https://www.w3.org/TR/SVG2/ I'm fairly certain I'd lose interest in getting it right before I got something better than most of those.

> The colors use traditional bicycle brown (#8B4513) and a classic blue for the pelican (#4169E1) with gold accents for the beak (#FFD700).

The output pelican is indeed blue. I can't fathom where the idea that this is "classic", or suitable for a pelican, could have come from.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#137
post #123
post #112

Earlier quoted context omitted.

It’s not that humans have perfect drawing skills, it’s that humans can judge their performance and get better over time. Ask 100 random people to draw a bike and in 10 minutes and they’ll on average suck while still beating the LLM’s here. Give em an incentive and 10 months and the average person is going to be able to make at least one quite decent drawing of a bike. The cost and speed advantage of LLM’s is real as…

> Ask 100 random people to draw a bike and in 10 minutes and they’ll on average suck while still beating the LLM’s here. Y'see, this is a prime example of what I meant with ""Average human" is a much lower bar than most people want to believe, mainly because most of us are average on most skills, and also overestimate our own competence". An expert artist can spend 10 minutes and end up with a brief sketch of a bike.…

As an objective criteria what percentage include peddles and a chain connecting one of the wheels? I quickly found a dozen and stopped counting. Now do the same for those LLM images and it’s clear humans win.

> ""Average human" is a much lower bar than most people want to believe

I have some basis for comparison. I’ve seen 6 years olds draw better bikes than those LLM’s.

Look through that list again the worst example does even have wheels, multiple of them have wheels without being connected to anything.

Now if you’re arguing the average human is worse than the average 6 year old I’m going to disagree here.

> Given mandatory art lessons in school are longer than 10 months, and yet those bike examples exist, I have no reason to believe this.

Art lessons don’t cumulatively spend 10 months teaching people how to draw a bike. I don’t think I cumulatively spent 6 months drawing anything. Painting, collage, sculpture, coloring, etc art covers a lot and wasn’t an every day or even every year thing. My mandatory collage class was art history, we didn’t create any art.

You may have spent more time in class studying drawing, but that’s not some universal average.

> If you automate it in literally the manner in this write-up (pairwise comparison via API calls to another model to get ELO ratings), ten thousand images is like $60-$90, which is on the low end for a human commission.

Not every one of those images had a price tag but one was 88 cents, * 10,000 = 8,800$ just to make the image for a test even at 4c/image your looking at 400$. Cheaper models existed but fairly consistently had worse performance.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#138

> most people find it difficult to remember the exact orientation of the frame. Isn't it Δ∇Λ welded together? The bottom left and right vertices are where the wheels are attached to, the middle bottom point is where the big gear with the pedals is. The lambda is for the front wheel because you wouldn't be able to turn it if it was attached to a delta. Right? I guess having my first bicycle be a cheap Soviet-era produ…

There are a lot of structural details that people tend to gloss over. This was illustrated by an Italian art project:

https://www.gianlucagimini.it/portfolio-item/velocipedia/

> back in 2009 I began pestering friends and random strangers. I would walk up to them with a pen and a sheet of paper asking that they immediately draw me a men’s bicycle, by heart. Soon I found out that when confronted with this odd request most people have a very hard time remembering exactly how a bike is made.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#139

As a control, he should go on fiver and have a human generate a pelican riding a bicycle, just to see what the eventual goal is.

Someone did this. Look at this sibling comment by ben_w https://news.ycombinator.com/item?id=44216284 about an old similar project.

> back in 2009 I began pestering friends and random strangers. I would walk up to them with a pen and a sheet of paper asking that they immediately draw me a men’s bicycle, by heart.

Someone commissioned to draw a bicycle on Fiverr would not have to rely on memory of what it should look like. It would take barely any time to just look up a reference.

Re: The last six months in LLMs, illustrated by pelicans on bicycles

#140
post #131

> If you lost interest in local models—like I did eight months ago—it’s worth paying attention to them again. They’ve got good now! > As a power user of these tools, I want to stay in complete control of what the inputs are. Features like ChatGPT memory are taking that control away from me. You reap what you sow.... > I already have a tool I built called shot-scraper, a CLI app that lets me take screenshots of web pa…

I had the LLM write a bash script for me that used my https://shot-scraper.datasette.io/ tool - on the basis that it was a neat opportunity to demonstrate another of my own projects. And honestly, even with LLM assistance getting Image Magick to output a 1200x600 image with two SVGs next to each other that are correctly resized to fill their half of the image sounds pretty tricky. Probably easier (for Claude) to achi…

> And honestly, even with LLM assistance getting Image Magick to output a 1200x600 image with two SVGs next to each other that are correctly resized to fill their half of the image sounds pretty tricky.

FWIW, the next project I want to look at after my current two, is a command-line tool to make this sort of thing easier. Likely featuring some sort of Lisp-like DSL to describe what to do with the input images.

Post reply on HN