I wish I could downvote this for Pelicanmaxxing lmao.
Are AI labs pelicanmaxxing?
31–40 of 253 posts
Re: Are AI labs pelicanmaxxing?
#32I feel like getting LLMs to spit out an SVG is akin to getting a human artist to draw something by just reciting a list of coordinates. It's insanely hard and unnatural. Image generation models nowadays can easily generate a photorealistic pelican riding a bicycle, where the bicycle has perfect structure. But it is, of course, only a raster image. It seems that we're missing a kind of step to decompose an image into…
Re: Are AI labs pelicanmaxxing?
#33This is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny . Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than a…
Re: Are AI labs pelicanmaxxing?
#34Test the LLLM against things you want it to do.
Asking questions that are absurd is like interviewing developers and asking absurd questions on the grounds that it tests creative and critical thinking.
Remember these Microsoft interview questions designed to identify the best developers?
"If you could eliminate one U.S. state, which one would it be?"
"How would you move Mount Fuji?"
Absurd interview questions have an air of legitimacy due to the quasi sophisticated justifications put forward for why they are good tests.
Absurd interview questions are not good tests of people or LLMs.
Relevant questions are good tests.
Re: Are AI labs pelicanmaxxing?
#35I feel like getting LLMs to spit out an SVG is akin to getting a human artist to draw something by just reciting a list of coordinates. It's insanely hard and unnatural. Image generation models nowadays can easily generate a photorealistic pelican riding a bicycle, where the bicycle has perfect structure. But it is, of course, only a raster image. It seems that we're missing a kind of step to decompose an image into…
The pelican on a bicycle test is specifically about generating an SVG, fyi, not a raster.
Re: Are AI labs pelicanmaxxing?
#36Oh great! You've now made it a lot easier for LLMs to train on this dataset! Your next iteration will need different animals and different transportation options. You'll run out after a few iterations.
Re: Are AI labs pelicanmaxxing?
#37Re: Are AI labs pelicanmaxxing?
#38https://playcode.io/blog/macbook-svg-benchmark I think we should stop using pelican benchmark.
I disagree with this in the blog post: > Every single one is a pelican, on a bicycle, first try. When every student gets an A, the exam has stopped grading. Numerous pelicans and their bikes are clearly horribly malformed. In fact none of the bike frames are correct. Fable and Opus come close, but the top of the diamond is disconnected in Fable's case and the head tube is misaligned with the front fork in Opus's case…
The problem isn't the test, its that is a public test.
Simon has previously said he has a list of secret prompts (at least one of which he "burned" as a demonstration a while ago). That's what makes it a good test - his commentary on the public test is something of a proxy for non-public tests. This makes it a good benchmark.
Re: Are AI labs pelicanmaxxing?
#39Looking for evidence of the same, but with another twist: checking if the models would choose to create a pelican on a bicycle, if no specific bird or method of transportation was specified.
My version of it: https://www.modelbias.ai/pelican-on-a-bicycle-test
Re: Are AI labs pelicanmaxxing?
#40So they aren't pelicanmaxxing but they are benchmaxxing in a way. The benefit of the pelican was originally that uplift on the pelican signaled an overall uplift on intelligence. I don't believe that is the case anymore and it is just another jagged edge of model intelligence.