I wish I could downvote this for Pelicanmaxxing lmao.
Would you prefer the term Pelicangate?
Are AI labs pelicanmaxxing?
131–140 of 253 posts
Re: Are AI labs pelicanmaxxing?
#132Re: Are AI labs pelicanmaxxing?
#133> Direction: All 21 pelican-bicycle images, across all seven labs, face right. No other animal/vehicle combination does that.
> However, facing right is common: 60% of all 1,008 images do it…
and the tables in “Evidence #5” to be anything but evidence the models have likely trained on pelican on bicycle data more than others.
The data clearly shows:
- 100% pelican on bicycle facing right
- significant skew to the right for bicycle-like vehicles
- significant preference for right facing for birds
Averaging those extreme results to “60%” to make it sound like it’s pretty fair because it’s close to “50%” isn’t statistically sound.
The methodology is generally unsound. There is no actual scoring with a well defined rubric, it’s just vibed with a single model (GPT 5.6 Luna).
The “not better at drawing” evidence are equally hard to take seriously when there is no clear, non-subjective indication of what better or worse is.
Re: Are AI labs pelicanmaxxing?
#134Re: Are AI labs pelicanmaxxing?
#135This is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny . Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than a…
What if they’re not pelicanmaxxing, but svgmaxxxing in general? Because otherwise using a LLM to generate complex svgs is pretty niche and what I thought made this a good benchmark when it was new - generalized programming and spatial knowledge. Obviously image gen in svg format is not a particularly hard problem if tackled directly on its own.
Re: Are AI labs pelicanmaxxing?
#136Earlier quoted context omitted.
But that's a genuine worthwhile capability. It's like benchmarkmaxxing on a weightlifting competition by getting really strong.
How useful actually is this? It generates SVGs of pelicans on bicycles, sure, and some of them are (almost) spatially correct. But, none of them look good . AI image generation suffers from this more generally. You can generate pictures of pelicans, sure. Newer models clearly generate images with more pelican-ness than before. But all of it is still uglier than sin. Drawing things accurately is one thing, making resu…
Re: Are AI labs pelicanmaxxing?
#137I recently had the following conversation with Claude: Me: how many P's are in the following text? [Pasted text] Claude: There are 14 P's, all lowercase (no capital P's) Me: how many in "strawberry"? Claude: there are 3 R's in the word "strawberry".
Re: Are AI labs pelicanmaxxing?
#138> All 21 pelican-bicycle images, across all seven labs, face right. No other animal/vehicle combination does that. > However, facing right is common: 60% of all 1,008 images do it. How common depends on the animal and the vehicle, and bicycles are one of the two vehicles where it’s strongest Of course the pelican on the bicycle is facing right. The drivetrain on a bicycle is on the right side. If you want any represe…
https://www.gianlucagimini.it/portfolio-item/velocipedia/
On the other hand, I notice that the prompt says to draw a pelican riding a bicycle, implying motion... and since most of us read left to right, I think it's usually natural to draw an object in motion moving left to right as well, which means the bicycle should be facing right. So maybe that specific setup is more natural here.
Either way, for humans bicycles are actually really hard to draw from memory. In fact, I substitute teach, and sometimes as an activity I have my students draw bicycles from memory in 60 seconds. Most make pretty serious errors, usually the frame or chain connections: they can tell it's wrong but still can't draw a more correct one. I use it as an object lesson about the difference between recognition and recall - most students never realize that much of their studying can end up being the former, when tests and life almost always ask for the latter. This helps explain why many students go from "that makes perfect sense" when going over review problems to a total mind blank only a few minutes later (especially in math!).
Re: Are AI labs pelicanmaxxing?
#139This is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny . Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than a…
What if they’re not pelicanmaxxing, but svgmaxxxing in general? Because otherwise using a LLM to generate complex svgs is pretty niche and what I thought made this a good benchmark when it was new - generalized programming and spatial knowledge. Obviously image gen in svg format is not a particularly hard problem if tackled directly on its own.
Mission fucking accomplished. https://xkcd.com/810/
Re: Are AI labs pelicanmaxxing?
#140This is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny . Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than a…