This is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny . Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than a…
What if they’re not pelicanmaxxing, but svgmaxxxing in general? Because otherwise using a LLM to generate complex svgs is pretty niche and what I thought made this a good benchmark when it was new - generalized programming and spatial knowledge. Obviously image gen in svg format is not a particularly hard problem if tackled directly on its own.
Are AI labs pelicanmaxxing?
241–250 of 253 posts
Re: Are AI labs pelicanmaxxing?
#242I'm curious if some of the animals / vehicules might force the models to use more tokens than others, and I could not find the token counts in the shared data, is it possible to publish it please? :)
Re: Are AI labs pelicanmaxxing?
#243Earlier quoted context omitted.
I’m waiting for the inevitable counterreaction to all this maxxing business, dubbed "minmaxxing"
In D&D, min-maxing (hyphen and one x) is already a used term for character optimization, minimizing undesirables and maximizing desirables (I guess). Minimizing peaks is probably not a good strategy generally. Even low peaks may have some benefit ("if you know your problem is in this domain and performance is critical, this tool offers a 3% advantage"). Minimizing troughs may be more attractive. People generally seem…
Re: Are AI labs pelicanmaxxing?
#244This is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny . Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than a…
Why is it hard to come up with tests that closely resemble real-world usecases where the models are being used?
We see the same problem in database benchmarking.
The good thing about the deliberately-non-real world pelican case is that it gives a general impression of how much the model is improving because it's not likely that it's being specifically targeted at it, rather than a 'real world case' which might have been specifically optimised for.
Re: Are AI labs pelicanmaxxing?
#245This is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny . Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than a…
Why is it hard to come up with tests that closely resemble real-world usecases where the models are being used?
They're mostly not very funny though.
Re: Are AI labs pelicanmaxxing?
#246Huh. They're not "Pelicanmaxxing"... they're Ottermaxxing. Take a look at the GLM 5.2 and Deepseek V4 "animal on a plane" examples. In every case, the animal is standing on top of the plane, a clear misunderstanding of the concept of "animal on a plane"... with the exception of the Otter. The otters are sitting in a seat on the plane, looking out the window. That's Ethan Mollick's "Otter On A Plane Using WiFi" image…
When LLMs use the pattern, they are often setting up a straw man and then knocking it over.
Re: Are AI labs pelicanmaxxing?
#247I'm glad someone ran the numbers on this. Every single Simon Willison post of an SVG is followed with someone dismissing it saying "I'm sure they train on it by now." This is despite a good blog post with sound logic on how easy that is to catch. [1] Glad to see someone took the time for a quantitative analysis of dumb little animals riding dumb little bikes. 1. https://simonwillison.net/2025/Nov/13/training-for-peli…
The very best svg pelican on a bile generation model. Just for laughs.
Re: Are AI labs pelicanmaxxing?
#248Earlier quoted context omitted.
The fear is that the SVGmaxxing is limited to "X doing Y". If such 'template maxxing' exists, it will break for other templates, e.g. "X not doing Y", "X and Y doing Z", "X doing Y doing Z", etc.
If a model can improve at drawing "X doing Y" and that prompt wasn't in the training set then it means it has improved its internal mapping from text-to-spatial-to-text.
Re: Are AI labs pelicanmaxxing?
#249Earlier quoted context omitted.
> I’m sure they will end up with a fun little toy from the whole endeavor they can play with for 2 weeks before abandoning. That's exactly the idea - except it's more like 2 hours before I prototype the next version. > Maybe the author will even feel bad about the carbon footprint of this whole exercise and buy some carbon offsets to make up for it. The passive aggressiveness of this is perfectly weighted and admirab…
I am glad you get to have fun while people you will never meet and live in Sweden or Virginia have to suffer higher electricity prices. But we all get to enjoy the warmer climate together. But at least you can enjoy your self having the plagiarizer build you a single use toy that only took the carbon footprint of entire nations to create.
Where I'm from data centers help the renewable mix by subsidizing transmission from other geographic zones. I'm actually improving the environment by using it.
[1]https://ourworldindata.org/how-much-energy-do-data-centers-a...
Re: Are AI labs pelicanmaxxing?
#250Earlier quoted context omitted.
ChatGPT 5.5 and Sol 5.6, Fable are good. I haven't really tested Opus 4.8, but 4.7 wasn't nearly as good as ChatGPT 5.5.
Interesting! The reason I asked is I've had poor results with Fable and 3d stuff. Its spatial awareness seems poor - doing things like rotating left and then right back, and then pitching nonsensically, just to try to capture a segment of a scene for a verification pass. And its placement often results in clipping, misrotations, and so on. It could well be that the exact domain matters more than the bigger picture co…
Mostly I have been using GPT 5.5 and now 5.6
That is notable because I do almost exclusively use Claude for coding.