Live data from Hacker News

Are AI labs pelicanmaxxing?

dylancastillo.co

171–180 of 253 posts

Re: Are AI labs pelicanmaxxing?

#171
post #84
post #59

Earlier quoted context omitted.

But that's a genuine worthwhile capability. It's like benchmarkmaxxing on a weightlifting competition by getting really strong.

https://www.youtube.com/watch?v=jgYYOUC10aM reminds me of this Key and Peele skit

Alternatively Mitchell and Webb https://www.youtube.com/watch?v=_pDTiFkXgEE

Re: Are AI labs pelicanmaxxing?

#173
post #59

Earlier quoted context omitted.

But that's a genuine worthwhile capability. It's like benchmarkmaxxing on a weightlifting competition by getting really strong.

Yes, but the original purpose of the benchmark (simonw, please correct me if I'm wrong!) was to test whether new models were good at novel problem solving. Things they haven't been trained on. So yes, getting better at generating SVGs is great news (and it seems they have been) but this particular benchmark still strikes me as largely worthless now, unless SVGs happen to be what you care about in particular when a ne…

I do have a private benchmark I use for that. I have a repository with an in-progress codebase and a very large planning document, I tell the new models to simply finish everything in the planning document, and I never push those changes. All of the recent models have essentially gotten it perfect which means that the models are all equivalently good for my level of needs.

Re: Are AI labs pelicanmaxxing?

#174
For an MVP I am building I asked to make a brain SVG and every attempt it has been doing it has been hillariously wrong and I had told Opus4.8/Fable to pick some SVG it could find online and it went ahead and still used its own thing that turned not too good. (full disclosure not particularly that good at front end stuff so relying heavily on Claude and Codex for it)

Re: Are AI labs pelicanmaxxing?

#175
post #44
post #17

This is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny . Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than a…

What if they’re not pelicanmaxxing, but svgmaxxxing in general? Because otherwise using a LLM to generate complex svgs is pretty niche and what I thought made this a good benchmark when it was new - generalized programming and spatial knowledge. Obviously image gen in svg format is not a particularly hard problem if tackled directly on its own.

Probably the case but not a bad one to "max"

Re: Are AI labs pelicanmaxxing?

#177
post #17

This is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny . Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than a…

It's all fun and games until that Pelican becomes self aware and asks themselves - why am I riding this bicycle ? I could be creating a Terminator 2 esque dystopian future myself.

And that's how we end up with PelicanSkynet

Re: Are AI labs pelicanmaxxing?

#178

i'm puzzled by the choice to have an llm judge the images.

I'm also puzzled by the fact that I expected EVERYONE to point this out but yours is the only comment about it.

Even if using LLM is the only way to do things at scale it does not mean that it's always the right tool.

Post reply on HN