Live data from Hacker News

Are AI labs pelicanmaxxing?

dylancastillo.co

121–130 of 253 posts

Re: Are AI labs pelicanmaxxing?

#121
As a unicyclist, all the sideways-riding caught my eye for being especially silly.

However, I'm very surprised that most of the models make the same sideways mistake with only some of the animals, and they do it consistently.

With most of the models, cat, raccoons, and otters are almost always riding sideways. Why is that?

Re: Are AI labs pelicanmaxxing?

#122
Huh. They're not "Pelicanmaxxing"... they're Ottermaxxing.

Take a look at the GLM 5.2 and Deepseek V4 "animal on a plane" examples. In every case, the animal is standing on top of the plane, a clear misunderstanding of the concept of "animal on a plane"... with the exception of the Otter. The otters are sitting in a seat on the plane, looking out the window.

That's Ethan Mollick's "Otter On A Plane Using WiFi" image benchmark.

https://www.oneusefulthing.org/p/the-recent-history-of-ai-in...

(Sometimes the Racoon is sitting inside the plane as well, but the racoon is a common backup benchmark. I'm surprised it wasn't also holding a sign saying that it loves trash.)

Also, Grok seemed to really really enjoy "whale on a plane" in that second round, and kudos to GPT Terra for deciding after 3 rounds that the user was terrible at spelling and generated "Antelope On A Plain".

EDIT: I promise I'm a human, but I did just notice my "that's not x... that's y" construction at the start. I am rather Claudepilled — my apologies.

Re: Are AI labs pelicanmaxxing?

#123

Earlier quoted context omitted.

well that holds IF svgmaxxing is 100% "code-writing-maxxing" ...which.. hmm I dunno if they are same or not

No, the point is that a general-ish ability to draw good SVGs is a useful ability in itself. People need SVGs for all sorts of purposes, and if AI can generate one for them, that's mostly useful (discussions about art and employment etc notwithstanding). That said, I think this would correlate relatively little with general programming ability. They're not unrelated, of course, but being able to generate code that pa…

If people need to generate good-ish SVG why not simply use a specialized model for a much better result and for far cheaper and quicker?

Why do LLMs need to be able to do this as well, but worse, slower and more expensive?

Re: Are AI labs pelicanmaxxing?

#124
post #17

This is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny . Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than a…

What if they’re just unsuccessful at it and something about a model grokking how to create a realistic visual representation of a pelican on a bicycle ends being key to the next OOM of capability unhobbling.

And since we have established this silly routine once, I must keep going and ask - yes, you have quite a collection of bona fide pelicans you’ve seen and photographed.

Have you physically seen all the other animals you’ve evaluated as well?

I was born in Soviet Russia, so trust but verify and if still around, maybe there pelican make LLM draw Simon ride bicycle and notice Python code improve.

Re: Are AI labs pelicanmaxxing?

#125
post #44
post #17

This is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny . Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than a…

What if they’re not pelicanmaxxing, but svgmaxxxing in general? Because otherwise using a LLM to generate complex svgs is pretty niche and what I thought made this a good benchmark when it was new - generalized programming and spatial knowledge. Obviously image gen in svg format is not a particularly hard problem if tackled directly on its own.

Not an expert on this, so won’t speculate regarding what traits would create robustness specifically attributed to SVG visual representation capability, but felt you might find this paper on reasoning models trained on physical world video data becoming better at general reasoning interesting:

https://arxiv.org/abs/2210.05359

Re: Are AI labs pelicanmaxxing?

#126

Earlier quoted context omitted.

No, the point is that a general-ish ability to draw good SVGs is a useful ability in itself. People need SVGs for all sorts of purposes, and if AI can generate one for them, that's mostly useful (discussions about art and employment etc notwithstanding). That said, I think this would correlate relatively little with general programming ability. They're not unrelated, of course, but being able to generate code that pa…

If people need to generate good-ish SVG why not simply use a specialized model for a much better result and for far cheaper and quicker? Why do LLMs need to be able to do this as well, but worse, slower and more expensive?

Agree directionally - even back in sonnet 3.5 days, I was helping some friends by showing them how to create intermediate representations for SVG building blocks mapping to parametrizable functions that can be used to make interactive SVG-rendered visualizations for various medical needs.

Re: Are AI labs pelicanmaxxing?

#127
post #47
post #17

This is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny . Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than a…

> Catching a lab cheating specifically on my one dumb benchmark would be really funny. Similar thing happened when TPC came up with SQL benchmarks. If you're not good at TPC, your engineering team is no good. If you're good at TPC, then (as a customer) we will actually include you in a bake-off benchmark for our specific problem. Winning on it is the price of admittance into the game, especially in a crowded market.…

Yea, I worked for a competitor to ATI back in the day and we were definitely Quakemaxxing. We were not so unethical as to try to detect the .EXE name so to put the GPU into a "cheating" mode only when that benchmark was run, but we did run Quake 3 pretty much constantly while trying to eke out a few more FPS...

Re: Are AI labs pelicanmaxxing?

#128
post #70
post #17

This is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny . Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than a…

Simon I hope from this day hence, your bio always includes: "Simon Willison, among other things, is an advocate for the inclusion of pelican geometry in LLM training datasets."

“Today, the NASDAQ fell by 18% moments after Simon Willison published his latest blog evaluating the Legend 7.4 model that rendered an animated SVG of a pelican on a bicycle and the pelican fell off”

Re: Are AI labs pelicanmaxxing?

#129
post #47

Earlier quoted context omitted.

> Catching a lab cheating specifically on my one dumb benchmark would be really funny. Similar thing happened when TPC came up with SQL benchmarks. If you're not good at TPC, your engineering team is no good. If you're good at TPC, then (as a customer) we will actually include you in a bake-off benchmark for our specific problem. Winning on it is the price of admittance into the game, especially in a crowded market.…

Yea, I worked for a competitor to ATI back in the day and we were definitely Quakemaxxing. We were not so unethical as to try to detect the .EXE name so to put the GPU into a "cheating" mode only when that benchmark was run, but we did run Quake 3 pretty much constantly while trying to eke out a few more FPS...

Wasn’t that almost the norm for both new OS drivers and subsequently GPU SoC firmware of ‘gamer cards’ to literally optimize for the latest AAA titles? I feel in that case, the interests aligned - in a world where I had a chance to actually play Crysis, I wouldn’t care why the frame rate was decent, no?

Re: Are AI labs pelicanmaxxing?

#130

Earlier quoted context omitted.

No, the point is that a general-ish ability to draw good SVGs is a useful ability in itself. People need SVGs for all sorts of purposes, and if AI can generate one for them, that's mostly useful (discussions about art and employment etc notwithstanding). That said, I think this would correlate relatively little with general programming ability. They're not unrelated, of course, but being able to generate code that pa…

If people need to generate good-ish SVG why not simply use a specialized model for a much better result and for far cheaper and quicker? Why do LLMs need to be able to do this as well, but worse, slower and more expensive?

Yeah, if you want to generate an aesthetically pleasing SVG, you'd be better off asking a pixel-based image-generation model for "vector art" and then deconstructing it into an equivalent SVG with something like LayerPeeler. https://layerpeeler.github.io/
Post reply on HN