Live data from Hacker News

Are AI labs pelicanmaxxing?

dylancastillo.co

191–200 of 253 posts

Re: Are AI labs pelicanmaxxing?

#191
post #59

Earlier quoted context omitted.

But that's a genuine worthwhile capability. It's like benchmarkmaxxing on a weightlifting competition by getting really strong.

> benchmarkmaxxing on a weightlifting competition If you stick to the benchpress, it's just "benchmaxxing".

I'm about ready for maxxmaxxing, where there's just maximally more of everything all the time

Re: Are AI labs pelicanmaxxing?

#192
post #44
post #17

This is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny . Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than a…

What if they’re not pelicanmaxxing, but svgmaxxxing in general? Because otherwise using a LLM to generate complex svgs is pretty niche and what I thought made this a good benchmark when it was new - generalized programming and spatial knowledge. Obviously image gen in svg format is not a particularly hard problem if tackled directly on its own.

They arguably might be, but I kind of want them to be!

Re: Are AI labs pelicanmaxxing?

#193

Earlier quoted context omitted.

> benchmarkmaxxing on a weightlifting competition If you stick to the benchpress, it's just "benchmaxxing".

I'm about ready for maxxmaxxing, where there's just maximally more of everything all the time

I’m waiting for the inevitable counterreaction to all this maxxing business, dubbed "minmaxxing"

Re: Are AI labs pelicanmaxxing?

#194

> All 21 pelican-bicycle images, across all seven labs, face right. No other animal/vehicle combination does that. > However, facing right is common: 60% of all 1,008 images do it. How common depends on the animal and the vehicle, and bicycles are one of the two vehicles where it’s strongest Of course the pelican on the bicycle is facing right. The drivetrain on a bicycle is on the right side. If you want any represe…

Perhaps japanese LLMs have a reason to do it differently:

https://www.cyclingweekly.com/news/japan-unveils-new-olympic...

Re: Are AI labs pelicanmaxxing?

#195
post #17

This is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny . Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than a…

I worked with image analytics years ago. In 2024 and cynical about LLM image capability, I tested different models to create "Two female octopi in a bar each having their own drink. One is wearing a yellow hat. The other is wearing a red hat. They have no human features." It went badly I think because the training visual training data associated with the word "female" was overwhelmingly biased in quantity and obviate…

Back in 2024 we didn’t really yet have image models that used bigger language models as their text encoders, so prompt understanding was still rather hit and miss, especially with negations involved.

Re: Are AI labs pelicanmaxxing?

#196
post #167

Earlier quoted context omitted.

Interestingly, there was an artist a few years back who (for an unrelated project) had almost 400 people across a range of ages draw a bicycle and 75% of those faced left! So this seems to actually go slightly against the human drawing intuition. https://www.gianlucagimini.it/portfolio-item/velocipedia/ On the other hand, I notice that the prompt says to draw a pelican riding a bicycle, implying motion... and since m…

That's interesting. I think the bikes-facing-left bias probably comes from how humans use bikes: the kickstand is on the left, so people likely hold and approach bikes from its left. Looking online, the kickstand is apparently on the left to avoid the gears. Based on the other comment about bike photography, it's interesting the same design choice makes humans and cameras/LLMs see bikes from different sides.

I think in general if you ask people to draw vehicles or animals (including humans) from the side, there will be this left-facing bias. Perhaps this is true only in left-to-right writing cultures, hard to say.

Re: Are AI labs pelicanmaxxing?

#197
post #153

Earlier quoted context omitted.

I use LLMs for 3D CAD design in OpenSCAD. There seems to be a very strong correlation between models that are good at SVG and models that are good at 3D CAD. Anecdote I know, but there does seem to be generalization going on here.

Which models have you found good for working with 3d stuff?

ChatGPT 5.5 and Sol 5.6, Fable are good.

I haven't really tested Opus 4.8, but 4.7 wasn't nearly as good as ChatGPT 5.5.

Re: Are AI labs pelicanmaxxing?

#198
There's every chance here I'm just being annoying pedant, but generating a bunch of SVGs of animals on vehicles doesn't mean that it's good at generating SVGs in general, just SVGs of animals on vehicles.

On the flip side, GPT 5.6 Sol did a pretty convincing render of a burglar eating salami.

I'd be curious to see how the other models on random things that are completely tangential to pelicans or bicycles.

Re: Are AI labs pelicanmaxxing?

#199
post #93

The pelican prompt is ridiculous. Test the LLLM against things you want it to do. Asking questions that are absurd is like interviewing developers and asking absurd questions on the grounds that it tests creative and critical thinking. Remember these Microsoft interview questions designed to identify the best developers? "If you could eliminate one U.S. state, which one would it be?" "How would you move Mount Fuji?"…

> Asking questions that are absurd is like interviewing developers and asking absurd questions on the grounds that it tests creative and critical thinking. Yes; what's wrong with that? Do you suppose that it doesn't test those qualities?

I, personally, think it does not. For the absurd task which has no testable solution, there is no grounded basis to judge creativity or critical thinking behind an answer. The only things that remain are "vibes" or "hilarity". Well, if these are what you want to test for, I do not disagree. But it isn't something I'd be interested in.

Re: Are AI labs pelicanmaxxing?

#200
post #44

Earlier quoted context omitted.

What if they’re not pelicanmaxxing, but svgmaxxxing in general? Because otherwise using a LLM to generate complex svgs is pretty niche and what I thought made this a good benchmark when it was new - generalized programming and spatial knowledge. Obviously image gen in svg format is not a particularly hard problem if tackled directly on its own.

I've been pretty happy with LLM svg based data plots I've asked, including log scaled axises and histograms. Definitely a first world problem of course.

But did they create those plots “by hand”, or did they use one of million open source libraries that create SVG plots? In my experience even SOTA models suck at generating SVG graphics like logos.
Post reply on HN