Earlier quoted context omitted.
And by a sample that has become increasingly known as a benchmark. Newer training data will contain more articles like this one, which naturally improves the capabilities of an LLM to estimate what’s considered a good „pelican on a bike“.
So what you really need to do is clone this blog post, find and replace pelican with any other noun, run all the tests, and publish that. Call it wikipediaslop.org
The last six months in LLMs, illustrated by pelicans on bicycles
121–130 of 244 posts
Re: The last six months in LLMs, illustrated by pelicans on bicycles
#122My only take home is they are all terrible and I should hire a professional.
Before that, you might ask ChatGPT to create a vector image of a pelican riding a bicycle and then running the output through a PNG to SVG converter... Result: https://www.dropbox.com/scl/fi/8b03yu5v58w0o5he1zayh/pelican... These are tough benchmarks to trial reasoning by having it _write_ an SVG file by hand and understanding how it's to be written to achieve this. Even a professional would struggle with that! It's…
Re: The last six months in LLMs, illustrated by pelicans on bicycles
#123Earlier quoted context omitted.
More that "these models work … like humans" (discretely or otherwise) does not imply the quotation. Most humans do not have perfect drawing skills and perfect knowledge about bikes and birds, they do not output such a simple drawing correctly 100% of the time. "Average human" is a much lower bar than most people want to believe, mainly because most of us are average on most skills, and also overestimate our own compe…
It’s not that humans have perfect drawing skills, it’s that humans can judge their performance and get better over time. Ask 100 random people to draw a bike and in 10 minutes and they’ll on average suck while still beating the LLM’s here. Give em an incentive and 10 months and the average person is going to be able to make at least one quite decent drawing of a bike. The cost and speed advantage of LLM’s is real as…
Y'see, this is a prime example of what I meant with ""Average human" is a much lower bar than most people want to believe, mainly because most of us are average on most skills, and also overestimate our own competence".
An expert artist can spend 10 minutes and end up with a brief sketch of a bike. You can witness this exact duration yourself (with non-bike examples) because of a challenge a few years back to draw the same picture in 10 minutes, 1 minute, and 10 seconds.
A normal person spending as much time as they like gets you the pictures that I linked to in the previous post, because they don't really know what a bike is. 45 examples of what normal people think a bike looks like: https://www.gianlucagimini.it/portfolio-item/velocipedia/
> Give em an incentive and 10 months and the average person is going to be able to make at least one quite decent drawing of a bike.
Given mandatory art lessons in school are longer than 10 months, and yet those bike examples exist, I have no reason to believe this.
> Ask a model for 10,000 drawings so you can pick the best and you get a marginal improvements based on random chance at a steep price.
If you do so as a human, rating and comparing images? Then the cost is your own time.
If you automate it in literally the manner in this write-up (pairwise comparison via API calls to another model to get ELO ratings), ten thousand images is like $60-$90, which is on the low end for a human commission.
Re: The last six months in LLMs, illustrated by pelicans on bicycles
#124"Say I have a wolf, a goat, and some cabbage, and I want to get them across a river. The wolf will eat the goat if they're left alone, which is bad. The goat will eat some cabbage, and will starve otherwise. How do I get them all across the river in the fewest trips?"
A child would pick up that you have plenty of cabbage, but can't leave the goat without it, lest it starve. Also, there's no mention of boat capacity, so you could just bring them all over at once. Useful? Sometimes. Intelligent? No.
Re: The last six months in LLMs, illustrated by pelicans on bicycles
#125Earlier quoted context omitted.
That’s ok, once bicycle “riding” pelicans become normative, we can ask it for images of pelicans humping bicycles. The number of subject-verb-objects are near infinite. All are imaginable, but most are not plausible. A plausibility machine (LLM) will struggle with the implausible, until it can abstract well.
> The number of subject-verb-objects are near infinite. All are imaginable, but most are not plausible Until there is enough unique/new subject-verb-objects examples/benchmarks so the trained model actually generalized it just like you did. (Public) Benchmarks needs to constantly evolve, otherwise they stop being useful.
Re: The last six months in LLMs, illustrated by pelicans on bicycles
#126It's not so great at bicycles, either. None of those are close to rideable. But bicycles are famously hard for artists as well. Cyclists can identify all of the parts, but if you don't ride a lot it can be surprisingly difficult to get all of the major bits of geometry right.
Re: The last six months in LLMs, illustrated by pelicans on bicycles
#127Earlier quoted context omitted.
Before that, you might ask ChatGPT to create a vector image of a pelican riding a bicycle and then running the output through a PNG to SVG converter... Result: https://www.dropbox.com/scl/fi/8b03yu5v58w0o5he1zayh/pelican... These are tough benchmarks to trial reasoning by having it _write_ an SVG file by hand and understanding how it's to be written to achieve this. Even a professional would struggle with that! It's…
I think you made an error there png is a bitmap format
Re: The last six months in LLMs, illustrated by pelicans on bicycles
#128> As a power user of these tools, I want to stay in complete control of what the inputs are. Features like ChatGPT memory are taking that control away from me.
You reap what you sow....
> I already have a tool I built called shot-scraper, a CLI app that lets me take screenshots of web pages and save them as images. I had Claude build me a web page that accepts ?left= and ?right= parameters pointing to image URLs and then embeds them side-by-side on a page. Then I could take screenshots of those two images side-by-side. I generated one of those for every possible match-up of my 34 pelican pictures—560 matches in total.
Surely it would have been easier to use a local tool like ImageMagick? You could even have the AI write a Bash script for you.
> ... but prompt injection is still a thing.
...Why wouldn't it always be? There's no quoting or escaping mechanism that's actually out-of-band.
> There’s this thing I’m calling the lethal trifecta, which is when you have an AI system that has access to private data, and potential exposure to malicious instructions—so other people can trick it into doing things... and there’s a mechanism to exfiltrate stuff.
People in 2025 actually need to be told this. Franklin missed the mark - people today will trip over themselves to give up both their security and their liberty for mere convenience.
Re: The last six months in LLMs, illustrated by pelicans on bicycles
#129Earlier quoted context omitted.
Yeah, this is the problem with benchmarks where the questions/problems are public. They're valuable for some months, until it bleeds into the training set. I'm certain a lot of the "improvements" we're seeing are just benchmarks leaking into the training set.
That’s ok, once bicycle “riding” pelicans become normative, we can ask it for images of pelicans humping bicycles. The number of subject-verb-objects are near infinite. All are imaginable, but most are not plausible. A plausibility machine (LLM) will struggle with the implausible, until it can abstract well.
Re: The last six months in LLMs, illustrated by pelicans on bicycles
#130Isn't it Δ∇Λ welded together? The bottom left and right vertices are where the wheels are attached to, the middle bottom point is where the big gear with the pedals is. The lambda is for the front wheel because you wouldn't be able to turn it if it was attached to a delta. Right?
I guess having my first bicycle be a cheap Soviet-era produced one paid off: I spent loads of time fidgeting with the chain tension, and pulling the chain back onto the gears, so I guess I had to stare at the frame way too much to forget even by today the way it looks.