Live data from Hacker News

Are AI labs pelicanmaxxing?

dylancastillo.co

41–50 of 253 posts

Re: Are AI labs pelicanmaxxing?

#41

The pelican prompt is ridiculous. Test the LLLM against things you want it to do. Asking questions that are absurd is like interviewing developers and asking absurd questions on the grounds that it tests creative and critical thinking. Remember these Microsoft interview questions designed to identify the best developers? "If you could eliminate one U.S. state, which one would it be?" "How would you move Mount Fuji?"…

> The pelican prompt is ridiculous

Yes, deliberately so.

It was never intended as a meaningful benchmark. The surprising thing was that for the first ~12 months performance on the stupid pelican benchmark did seem to correspond to the performance of the models on other tasks.

That pattern no longer holds - Fable 5 and GPT-5.6 have both been out-pelicaned by lesser models now.

Re: Are AI labs pelicanmaxxing?

#42

You seem to assume training on pelican would not result in improved performance on other similar tasks. Why?

He didn't. That's why the article exists. You have to do the science to see if it does.

He was asking the question - do we see gains across other tasks? The underlying question was: Is the additional attention given to this specific task creating a false impression of progress?

Re: Are AI labs pelicanmaxxing?

#43

The pelican prompt is ridiculous. Test the LLLM against things you want it to do. Asking questions that are absurd is like interviewing developers and asking absurd questions on the grounds that it tests creative and critical thinking. Remember these Microsoft interview questions designed to identify the best developers? "If you could eliminate one U.S. state, which one would it be?" "How would you move Mount Fuji?"…

As I understand it; the point is to ask for an SVG which would demonstrate a conceptual understanding of what is being asked for and that is an important test IMO.

What sufficiently hard, but useful, problem would you ask the model for?

Re: Are AI labs pelicanmaxxing?

#44
post #17

This is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny . Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than a…

What if they’re not pelicanmaxxing, but svgmaxxxing in general?

Because otherwise using a LLM to generate complex svgs is pretty niche and what I thought made this a good benchmark when it was new - generalized programming and spatial knowledge.

Obviously image gen in svg format is not a particularly hard problem if tackled directly on its own.

Re: Are AI labs pelicanmaxxing?

#45
post #16

> The more plausible story is SVGmaxxing Exactly--and you have to ask yourself at this point what "maxxing" really means, since "get better at drawing SVGs" is a useful skill.

Really awful how the AI labs are skillmaxxing /s Pelicans aside, we need to remember that benchmarks are the only good quantitative way we have of comparing models. If someone has complaints about “benchmaxxing”, please ask them to contribute a better benchmark! It is valuable work and very appreciated.

> If someone has complaints about “benchmaxxing”, please ask them to contribute a better benchmark!

I don't think I'd go that far!

When someone says a model has been benchmaxxed, what they really mean is that it performs better in benchmarks compared to their real world experience. That's a real thing, I've certainly experienced it with some models.

...my take is that some things in life just resist quantitative measurements. Who is the best job candidate? What is the best programming language? Add AI models to the pile.

Re: Are AI labs pelicanmaxxing?

#46
https://en.wikipedia.org/wiki/Goodhart%27s_law

"Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes."

Or the more pop layman version

"When a measure becomes a metric/KPI, it ceases to be a good measure."

Story time, I live in Argentina, and we don't have Big Macs, the main Mc Donald's brand, here, because during the CFK presidency, one of her tactics was to Goodhart economic metrics. Even the informal obscure ones like the [Big Mac Index](https://en.wikipedia.org/wiki/Big_Mac_Index), I don't know the precise details, but the Big Mac ended up being a very cheap item, like 2 or 3 times cheaper than actual menu items, but it was never on the advertised menu, and it also ended up being very small compared to the other burgers, so it wasn't even like a hack, a shrinkflation type of deal.

But hey, anyone who read the Big Mac Index table would never find Argentina at the bottom of that list along with a couple of other countries with bad brands, so the ploy worked. And now we live with the aftershock, the brand never really turned around, other brands with ridiculous names took over it like the McTasty, which makes me sound like that skit from Tarantino's Pulp Fiction.

Re: Are AI labs pelicanmaxxing?

#47
post #17

This is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny . Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than a…

> Catching a lab cheating specifically on my one dumb benchmark would be really funny.

Similar thing happened when TPC came up with SQL benchmarks.

If you're not good at TPC, your engineering team is no good.

If you're good at TPC, then (as a customer) we will actually include you in a bake-off benchmark for our specific problem.

Winning on it is the price of admittance into the game, especially in a crowded market.

But how narrowly you benchmarket matters, you can't just hard-code that specific scenario & not fix anything adjacent while you're at it.

For example when it comes to GPUs, the "Quack3" (sic) benchmark on ATI cards comes to mind.

Re: Are AI labs pelicanmaxxing?

#48
post #44
post #17

This is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny . Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than a…

What if they’re not pelicanmaxxing, but svgmaxxxing in general? Because otherwise using a LLM to generate complex svgs is pretty niche and what I thought made this a good benchmark when it was new - generalized programming and spatial knowledge. Obviously image gen in svg format is not a particularly hard problem if tackled directly on its own.

I agree, other formats, both textual and binary should be tested.

Re: Are AI labs pelicanmaxxing?

#49
I am not sure if this is how it works but let's say there was a reddit thread talking about the pelican benchmark and in it someone posts mockup examples of what an ideal result would look like

aren't some LLM going to digest that thread at some point and indirectly learn from it?

basically my point is originally this was a good benchmark because it was an absurd never-seen-before thing, but now that it is in content, some models are going to get a benefit in education?

you'd need the "AI" equivalent of an old-school "google whack", something with no previous results

* https://en.wikipedia.org/wiki/Googlewhack

Re: Are AI labs pelicanmaxxing?

#50
post #44
post #17

This is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny . Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than a…

What if they’re not pelicanmaxxing, but svgmaxxxing in general? Because otherwise using a LLM to generate complex svgs is pretty niche and what I thought made this a good benchmark when it was new - generalized programming and spatial knowledge. Obviously image gen in svg format is not a particularly hard problem if tackled directly on its own.

Addressed in the article, in case you're curious.
Post reply on HN