Live data from Hacker News

Are AI labs pelicanmaxxing?

dylancastillo.co

181–190 of 253 posts

Re: Are AI labs pelicanmaxxing?

#181
post #128
post #70

Earlier quoted context omitted.

Simon I hope from this day hence, your bio always includes: "Simon Willison, among other things, is an advocate for the inclusion of pelican geometry in LLM training datasets."

“Today, the NASDAQ fell by 18% moments after Simon Willison published his latest blog evaluating the Legend 7.4 model that rendered an animated SVG of a pelican on a bicycle and the pelican fell off”

When asked for comment an anthropic spokesperson said "it's very unusual for the pelican to fall off. They are built to very rigorous SVG engineering standards"

Re: Are AI labs pelicanmaxxing?

#182
post #169

Earlier quoted context omitted.

I'd be very impressed by anyone who can produce that in Figma or Word in ten minutes.

These are rectangles with some fill and strokes, textboxes lines and arrows. If the text has already been written, and the author has a clear idea of the flow all it takes are some a dozen or so points and click, and setting a couple of fills, then copy pastes, and finally some resizes, selects and moves. This design is trivial. I admit it would be hard to achieve in Word (or at least for me because I don’t know how…

> I’m sure they will end up with a fun little toy from the whole endeavor they can play with for 2 weeks before abandoning.

That's exactly the idea - except it's more like 2 hours before I prototype the next version.

> Maybe the author will even feel bad about the carbon footprint of this whole exercise and buy some carbon offsets to make up for it.

The passive aggressiveness of this is perfectly weighted and admirably phrased.

Re: Are AI labs pelicanmaxxing?

#184
Okay, first off, without honestly reading the whole article (I tried but I just don’t have the patience), a quick red flag is the final analysis is only through Fable? Isn’t that inherently going to introduce unwanted bias?

Whatever. Doesn’t really matter much. My next thought is, I get that requiring an SVG is adding an extra layer of complexity as far as the art goes, but why is nobody talking about that the actual art is absolute trash?

I get it. It’s basically a meme at this point and it’s a fun game to play with the models. But my thought is it should be illuminating to anyone who is an artist that LLMs are still a long way off from taking your job :)

Re: Are AI labs pelicanmaxxing?

#186
post #44
post #17

This is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny . Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than a…

What if they’re not pelicanmaxxing, but svgmaxxxing in general? Because otherwise using a LLM to generate complex svgs is pretty niche and what I thought made this a good benchmark when it was new - generalized programming and spatial knowledge. Obviously image gen in svg format is not a particularly hard problem if tackled directly on its own.

Generating images not by pixels but by code and instructions is niche to you?

Re: Are AI labs pelicanmaxxing?

#187
I would expect so.

In the last century, as spec tests for C and C++ compilers, databases, Java application servers became trendy, all vendors were optimising for great articles on the respective technical magazines.

Re: Are AI labs pelicanmaxxing?

#189

Earlier quoted context omitted.

How useful actually is this? It generates SVGs of pelicans on bicycles, sure, and some of them are (almost) spatially correct. But, none of them look good . AI image generation suffers from this more generally. You can generate pictures of pelicans, sure. Newer models clearly generate images with more pelican-ness than before. But all of it is still uglier than sin. Drawing things accurately is one thing, making resu…

Just today I asked Claude Fable to take an SVG map and add a circle with a 150 mile radius around a specific city, cutting off the circle at the edge of certain boundaries. It came back a minute later with more or less exactly what I wanted, saving me maybe 5 to 10 minutes of photoshop time. It’s almost certainly not perfect, but I didn’t need it to be perfect, just a legible representation for a group of 20ish peopl…

OTOH I had Fable laying out a 3d broadcast room, and as part of the assets it sourced, it grabbed a camera. Makes sense. There were a lot of problems with the scene - backwards facing props and clipping everywhere. But the true coup de grace was a mansized, more like a giant sized, 35mm camera pointing at everything. I was seriously confused for a minute until I realized what had happened.

Re: Are AI labs pelicanmaxxing?

#190
post #153

Earlier quoted context omitted.

The fear is that the SVGmaxxing is limited to "X doing Y". If such 'template maxxing' exists, it will break for other templates, e.g. "X not doing Y", "X and Y doing Z", "X doing Y doing Z", etc.

I use LLMs for 3D CAD design in OpenSCAD. There seems to be a very strong correlation between models that are good at SVG and models that are good at 3D CAD. Anecdote I know, but there does seem to be generalization going on here.

Which models have you found good for working with 3d stuff?
Post reply on HN