Live data from Hacker News

Are AI labs pelicanmaxxing?

dylancastillo.co

241–250 of 253 posts

Re: Are AI labs pelicanmaxxing?

#241
post #44
post #17

This is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny . Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than a…

What if they’re not pelicanmaxxing, but svgmaxxxing in general? Because otherwise using a LLM to generate complex svgs is pretty niche and what I thought made this a good benchmark when it was new - generalized programming and spatial knowledge. Obviously image gen in svg format is not a particularly hard problem if tackled directly on its own.

svgmaxxing, please do it!

https://xkcd.com/810/

Re: Are AI labs pelicanmaxxing?

#242
Nice article! Thanks for experimenting and writing it.

I'm curious if some of the animals / vehicules might force the models to use more tokens than others, and I could not find the token counts in the shared data, is it possible to publish it please? :)

Re: Are AI labs pelicanmaxxing?

#243

Earlier quoted context omitted.

I’m waiting for the inevitable counterreaction to all this maxxing business, dubbed "minmaxxing"

In D&D, min-maxing (hyphen and one x) is already a used term for character optimization, minimizing undesirables and maximizing desirables (I guess). Minimizing peaks is probably not a good strategy generally. Even low peaks may have some benefit ("if you know your problem is in this domain and performance is critical, this tool offers a 3% advantage"). Minimizing troughs may be more attractive. People generally seem…

Yes, it was a subtle referencing to rpg minmaxing, and also the minimax algorithm, beyond the more surface-level idea that "maximalist" trends tend to be followed by "minimalist" ones as a natural counterreaction.

Re: Are AI labs pelicanmaxxing?

#244
post #220
post #17

This is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny . Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than a…

Why is it hard to come up with tests that closely resemble real-world usecases where the models are being used?

Would it be any more applicable?

We see the same problem in database benchmarking.

The good thing about the deliberately-non-real world pelican case is that it gives a general impression of how much the model is improving because it's not likely that it's being specifically targeted at it, rather than a 'real world case' which might have been specifically optimised for.

Re: Are AI labs pelicanmaxxing?

#245
post #220
post #17

This is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny . Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than a…

Why is it hard to come up with tests that closely resemble real-world usecases where the models are being used?

It's not that hard - there are plenty of those.

They're mostly not very funny though.

Re: Are AI labs pelicanmaxxing?

#246

Huh. They're not "Pelicanmaxxing"... they're Ottermaxxing. Take a look at the GLM 5.2 and Deepseek V4 "animal on a plane" examples. In every case, the animal is standing on top of the plane, a clear misunderstanding of the concept of "animal on a plane"... with the exception of the Otter. The otters are sitting in a seat on the plane, looking out the window. That's Ethan Mollick's "Otter On A Plane Using WiFi" image…

To be fair, when I see that pattern after a relevant question, it isn't much of an AI tell. Consider this exchange: A:"Are you hungry?" B:"I'm not hungry, I'm thirsty."

When LLMs use the pattern, they are often setting up a straw man and then knocking it over.

Re: Are AI labs pelicanmaxxing?

#247

I'm glad someone ran the numbers on this. Every single Simon Willison post of an SVG is followed with someone dismissing it saying "I'm sure they train on it by now." This is despite a good blog post with sound logic on how easy that is to catch. [1] Glad to see someone took the time for a quantitative analysis of dumb little animals riding dumb little bikes. 1. https://simonwillison.net/2025/Nov/13/training-for-peli…

If I were running one of these mega AI companies I would set aside a tiny team to produce and release a pelican model, explicitly trained for this.

The very best svg pelican on a bile generation model. Just for laughs.

Re: Are AI labs pelicanmaxxing?

#248

Earlier quoted context omitted.

The fear is that the SVGmaxxing is limited to "X doing Y". If such 'template maxxing' exists, it will break for other templates, e.g. "X not doing Y", "X and Y doing Z", "X doing Y doing Z", etc.

If a model can improve at drawing "X doing Y" and that prompt wasn't in the training set then it means it has improved its internal mapping from text-to-spatial-to-text.

The improvement would be limited to the template "X doing Y". It would not be a general improvement unless hundreds or thousands of diverse templates were used.

Re: Are AI labs pelicanmaxxing?

#249
post #182

Earlier quoted context omitted.

> I’m sure they will end up with a fun little toy from the whole endeavor they can play with for 2 weeks before abandoning. That's exactly the idea - except it's more like 2 hours before I prototype the next version. > Maybe the author will even feel bad about the carbon footprint of this whole exercise and buy some carbon offsets to make up for it. The passive aggressiveness of this is perfectly weighted and admirab…

I am glad you get to have fun while people you will never meet and live in Sweden or Virginia have to suffer higher electricity prices. But we all get to enjoy the warmer climate together. But at least you can enjoy your self having the plagiarizer build you a single use toy that only took the carbon footprint of entire nations to create.

AI consumed around 0.5% of the world’s electricity in 2025[1]

Where I'm from data centers help the renewable mix by subsidizing transmission from other geographic zones. I'm actually improving the environment by using it.

[1]https://ourworldindata.org/how-much-energy-do-data-centers-a...

Re: Are AI labs pelicanmaxxing?

#250
post #197

Earlier quoted context omitted.

ChatGPT 5.5 and Sol 5.6, Fable are good. I haven't really tested Opus 4.8, but 4.7 wasn't nearly as good as ChatGPT 5.5.

Interesting! The reason I asked is I've had poor results with Fable and 3d stuff. Its spatial awareness seems poor - doing things like rotating left and then right back, and then pitching nonsensically, just to try to capture a segment of a scene for a verification pass. And its placement often results in clipping, misrotations, and so on. It could well be that the exact domain matters more than the bigger picture co…

That's interesting. I only did a couple of attempts with Fable and it seemed fine.

Mostly I have been using GPT 5.5 and now 5.6

That is notable because I do almost exclusively use Claude for coding.

Post reply on HN