Live data from Hacker News

Are AI labs pelicanmaxxing?

dylancastillo.co

31–40 of 253 posts

Re: Are AI labs pelicanmaxxing?

#32
post #24

I feel like getting LLMs to spit out an SVG is akin to getting a human artist to draw something by just reciting a list of coordinates. It's insanely hard and unnatural. Image generation models nowadays can easily generate a photorealistic pelican riding a bicycle, where the bicycle has perfect structure. But it is, of course, only a raster image. It seems that we're missing a kind of step to decompose an image into…

Image models that support text output like Image2, or general text models that can read images like Claude can vectorize raster images. But they aren't very good at it, doing it manually in Inkscape still produces better quality even when done by non-artists.

Re: Are AI labs pelicanmaxxing?

#33
post #17

This is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny . Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than a…

[dead]

Re: Are AI labs pelicanmaxxing?

#34
The pelican prompt is ridiculous.

Test the LLLM against things you want it to do.

Asking questions that are absurd is like interviewing developers and asking absurd questions on the grounds that it tests creative and critical thinking.

Remember these Microsoft interview questions designed to identify the best developers?

"If you could eliminate one U.S. state, which one would it be?"

"How would you move Mount Fuji?"

Absurd interview questions have an air of legitimacy due to the quasi sophisticated justifications put forward for why they are good tests.

Absurd interview questions are not good tests of people or LLMs.

Relevant questions are good tests.

Re: Are AI labs pelicanmaxxing?

#35
post #24

I feel like getting LLMs to spit out an SVG is akin to getting a human artist to draw something by just reciting a list of coordinates. It's insanely hard and unnatural. Image generation models nowadays can easily generate a photorealistic pelican riding a bicycle, where the bicycle has perfect structure. But it is, of course, only a raster image. It seems that we're missing a kind of step to decompose an image into…

The pelican on a bicycle test is specifically about generating an SVG, fyi, not a raster.

I know. I'm just thinking about how to make AI create SVGs better... in theory, a sufficiently smart AI could "generate an image in its head", think about it, and then output the SVG paths to produce said image. Intuitively that would be somewhat closer to how human artists convert artistic visions into a sequence of arm movements while holding a brush (obviously, humans don't hold a fully formed, photorealistic image in the head while drawing, but rather vague concepts, but still).

Re: Are AI labs pelicanmaxxing?

#36
post #30

Oh great! You've now made it a lot easier for LLMs to train on this dataset! Your next iteration will need different animals and different transportation options. You'll run out after a few iterations.

"benchmaxxing by generalizing" is not really benchmaxxing

Re: Are AI labs pelicanmaxxing?

#38
post #28
post #9

https://playcode.io/blog/macbook-svg-benchmark I think we should stop using pelican benchmark.

I disagree with this in the blog post: > Every single one is a pelican, on a bicycle, first try. When every student gets an A, the exam has stopped grading. Numerous pelicans and their bikes are clearly horribly malformed. In fact none of the bike frames are correct. Fable and Opus come close, but the top of the diamond is disconnected in Fable's case and the head tube is misaligned with the front fork in Opus's case…

Agreed. And more; the Macbooks are pretty much the same - some are god approximations, some are terrible, all of them are recognisably a MacBook. And if you start using it they can train on it.

The problem isn't the test, its that is a public test.

Simon has previously said he has a list of secret prompts (at least one of which he "burned" as a demonstration a while ago). That's what makes it a good test - his commentary on the public test is something of a proxy for non-public tests. This makes it a good benchmark.

Re: Are AI labs pelicanmaxxing?

#39
This is funny, I actually did a similar experiment just yesterday.

Looking for evidence of the same, but with another twist: checking if the models would choose to create a pelican on a bicycle, if no specific bird or method of transportation was specified.

My version of it: https://www.modelbias.ai/pelican-on-a-bicycle-test

Re: Are AI labs pelicanmaxxing?

#40
I've had the feeling that labs aren't pelicanmaxxing specifically but that they do have some sort of RL environment for SVGs that they are letting the AIs overcook in. Specifically I'm thinking of the gemini 3.1 pro annoucnement that seemed to have a huge leap in animated SVG performance but not much else impressive about it.

So they aren't pelicanmaxxing but they are benchmaxxing in a way. The benefit of the pelican was originally that uplift on the pelican signaled an overall uplift on intelligence. I don't believe that is the case anymore and it is just another jagged edge of model intelligence.

Post reply on HN