Earlier quoted context omitted.
But that's a genuine worthwhile capability. It's like benchmarkmaxxing on a weightlifting competition by getting really strong.
> benchmarkmaxxing on a weightlifting competition If you stick to the benchpress, it's just "benchmaxxing".
Are AI labs pelicanmaxxing?
191–200 of 253 posts
Re: Are AI labs pelicanmaxxing?
#192This is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny . Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than a…
What if they’re not pelicanmaxxing, but svgmaxxxing in general? Because otherwise using a LLM to generate complex svgs is pretty niche and what I thought made this a good benchmark when it was new - generalized programming and spatial knowledge. Obviously image gen in svg format is not a particularly hard problem if tackled directly on its own.
Re: Are AI labs pelicanmaxxing?
#193Earlier quoted context omitted.
> benchmarkmaxxing on a weightlifting competition If you stick to the benchpress, it's just "benchmaxxing".
I'm about ready for maxxmaxxing, where there's just maximally more of everything all the time
Re: Are AI labs pelicanmaxxing?
#194> All 21 pelican-bicycle images, across all seven labs, face right. No other animal/vehicle combination does that. > However, facing right is common: 60% of all 1,008 images do it. How common depends on the animal and the vehicle, and bicycles are one of the two vehicles where it’s strongest Of course the pelican on the bicycle is facing right. The drivetrain on a bicycle is on the right side. If you want any represe…
https://www.cyclingweekly.com/news/japan-unveils-new-olympic...
Re: Are AI labs pelicanmaxxing?
#195This is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny . Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than a…
I worked with image analytics years ago. In 2024 and cynical about LLM image capability, I tested different models to create "Two female octopi in a bar each having their own drink. One is wearing a yellow hat. The other is wearing a red hat. They have no human features." It went badly I think because the training visual training data associated with the word "female" was overwhelmingly biased in quantity and obviate…
Re: Are AI labs pelicanmaxxing?
#196Earlier quoted context omitted.
Interestingly, there was an artist a few years back who (for an unrelated project) had almost 400 people across a range of ages draw a bicycle and 75% of those faced left! So this seems to actually go slightly against the human drawing intuition. https://www.gianlucagimini.it/portfolio-item/velocipedia/ On the other hand, I notice that the prompt says to draw a pelican riding a bicycle, implying motion... and since m…
That's interesting. I think the bikes-facing-left bias probably comes from how humans use bikes: the kickstand is on the left, so people likely hold and approach bikes from its left. Looking online, the kickstand is apparently on the left to avoid the gears. Based on the other comment about bike photography, it's interesting the same design choice makes humans and cameras/LLMs see bikes from different sides.
Re: Are AI labs pelicanmaxxing?
#197Earlier quoted context omitted.
I use LLMs for 3D CAD design in OpenSCAD. There seems to be a very strong correlation between models that are good at SVG and models that are good at 3D CAD. Anecdote I know, but there does seem to be generalization going on here.
Which models have you found good for working with 3d stuff?
I haven't really tested Opus 4.8, but 4.7 wasn't nearly as good as ChatGPT 5.5.
Re: Are AI labs pelicanmaxxing?
#198On the flip side, GPT 5.6 Sol did a pretty convincing render of a burglar eating salami.
I'd be curious to see how the other models on random things that are completely tangential to pelicans or bicycles.
Re: Are AI labs pelicanmaxxing?
#199The pelican prompt is ridiculous. Test the LLLM against things you want it to do. Asking questions that are absurd is like interviewing developers and asking absurd questions on the grounds that it tests creative and critical thinking. Remember these Microsoft interview questions designed to identify the best developers? "If you could eliminate one U.S. state, which one would it be?" "How would you move Mount Fuji?"…
> Asking questions that are absurd is like interviewing developers and asking absurd questions on the grounds that it tests creative and critical thinking. Yes; what's wrong with that? Do you suppose that it doesn't test those qualities?
Re: Are AI labs pelicanmaxxing?
#200Earlier quoted context omitted.
What if they’re not pelicanmaxxing, but svgmaxxxing in general? Because otherwise using a LLM to generate complex svgs is pretty niche and what I thought made this a good benchmark when it was new - generalized programming and spatial knowledge. Obviously image gen in svg format is not a particularly hard problem if tackled directly on its own.
I've been pretty happy with LLM svg based data plots I've asked, including log scaled axises and histograms. Definitely a first world problem of course.