I find it humorous that the animal + plane combo appears to be such an outlier. I assume this is due to the models assuming the user mean plain and misspelled it in the prompt.
GLM has 2 combos of "on a plane" literally sitting inside a plane, with a window and a bit of wing showing. That's funny.
Are AI labs pelicanmaxxing?
91–100 of 253 posts
Re: Are AI labs pelicanmaxxing?
#92Re: Are AI labs pelicanmaxxing?
#93The pelican prompt is ridiculous. Test the LLLM against things you want it to do. Asking questions that are absurd is like interviewing developers and asking absurd questions on the grounds that it tests creative and critical thinking. Remember these Microsoft interview questions designed to identify the best developers? "If you could eliminate one U.S. state, which one would it be?" "How would you move Mount Fuji?"…
Yes; what's wrong with that?
Do you suppose that it doesn't test those qualities?
Re: Are AI labs pelicanmaxxing?
#94https://en.wikipedia.org/wiki/Goodhart%27s_law "Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes." Or the more pop layman version "When a measure becomes a metric/KPI, it ceases to be a good measure." Story time, I live in Argentina, and we don't have Big Macs, the main Mc Donald's brand, here, because during the CFK presidency, one of her tactics was to G…
Re: Are AI labs pelicanmaxxing?
#95This is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny . Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than a…
What if they’re not pelicanmaxxing, but svgmaxxxing in general? Because otherwise using a LLM to generate complex svgs is pretty niche and what I thought made this a good benchmark when it was new - generalized programming and spatial knowledge. Obviously image gen in svg format is not a particularly hard problem if tackled directly on its own.
Re: Are AI labs pelicanmaxxing?
#96Re: Are AI labs pelicanmaxxing?
#97Earlier quoted context omitted.
But that's a genuine worthwhile capability. It's like benchmarkmaxxing on a weightlifting competition by getting really strong.
How useful actually is this? It generates SVGs of pelicans on bicycles, sure, and some of them are (almost) spatially correct. But, none of them look good . AI image generation suffers from this more generally. You can generate pictures of pelicans, sure. Newer models clearly generate images with more pelican-ness than before. But all of it is still uglier than sin. Drawing things accurately is one thing, making resu…
Re: Are AI labs pelicanmaxxing?
#98Earlier quoted context omitted.
But that's a genuine worthwhile capability. It's like benchmarkmaxxing on a weightlifting competition by getting really strong.
How useful actually is this? It generates SVGs of pelicans on bicycles, sure, and some of them are (almost) spatially correct. But, none of them look good . AI image generation suffers from this more generally. You can generate pictures of pelicans, sure. Newer models clearly generate images with more pelican-ness than before. But all of it is still uglier than sin. Drawing things accurately is one thing, making resu…
Re: Are AI labs pelicanmaxxing?
#99Tomato, tomato
Re: Are AI labs pelicanmaxxing?
#100This is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny . Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than a…
What if they’re not pelicanmaxxing, but svgmaxxxing in general? Because otherwise using a LLM to generate complex svgs is pretty niche and what I thought made this a good benchmark when it was new - generalized programming and spatial knowledge. Obviously image gen in svg format is not a particularly hard problem if tackled directly on its own.