Live data from Hacker News

Are AI labs pelicanmaxxing?

dylancastillo.co

161–170 of 253 posts

Re: Are AI labs pelicanmaxxing?

#161

Earlier quoted context omitted.

Yes, but the original purpose of the benchmark (simonw, please correct me if I'm wrong!) was to test whether new models were good at novel problem solving. Things they haven't been trained on. So yes, getting better at generating SVGs is great news (and it seems they have been) but this particular benchmark still strikes me as largely worthless now, unless SVGs happen to be what you care about in particular when a ne…

I've mostly been running LLMs on my own hardware so the phrase "been at it for about 20 minutes" gave me pause. My first instinct was to ask "but on what hardware?" but I suppose one advantage of closed models is that there's a somewhat-consistent cognitive-effort-to-time ratio.

Fair, I don’t know if this helps but it was Opus 4.8 (high effort) via Claude Desktop, but through my org’s LLM gateway so hard to compare really. But it’s usually pretty quick.

(don’t tell my boss.)

Re: Are AI labs pelicanmaxxing?

#162
I feel like the simplest Pelicanmaxxing method is just to teach the model that whenever a user asks for a svg illustration of something with no other clarification, it should default to making it as detailed and pretty as possible. This would make every single svg from that model look better and not just the pelican on a bicycle

Re: Are AI labs pelicanmaxxing?

#163
post #154

Earlier quoted context omitted.

If people need to generate good-ish SVG why not simply use a specialized model for a much better result and for far cheaper and quicker? Why do LLMs need to be able to do this as well, but worse, slower and more expensive?

For example I do SVG diagrams to explain architecture and specific data flows in a multi-tier app. LLMs do a great job because they understand both code and SVGs well. Edit: An example for a synth I'm buulding: https://imgur.com/a/U694Ek7 The irony of having to post it as a PNG isn't lost on me...

This is not a strong use case. I could have made this in Microsoft word when I was 13 years old.

Why on earth would you let an LLM do this when it would take you 10 min to do this in Figma or Inkscape or even just Word

Re: Are AI labs pelicanmaxxing?

#164
post #44
post #17

This is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny . Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than a…

What if they’re not pelicanmaxxing, but svgmaxxxing in general? Because otherwise using a LLM to generate complex svgs is pretty niche and what I thought made this a good benchmark when it was new - generalized programming and spatial knowledge. Obviously image gen in svg format is not a particularly hard problem if tackled directly on its own.

I've been pretty happy with LLM svg based data plots I've asked, including log scaled axises and histograms. Definitely a first world problem of course.

Re: Are AI labs pelicanmaxxing?

#165
post #154

Earlier quoted context omitted.

If people need to generate good-ish SVG why not simply use a specialized model for a much better result and for far cheaper and quicker? Why do LLMs need to be able to do this as well, but worse, slower and more expensive?

For example I do SVG diagrams to explain architecture and specific data flows in a multi-tier app. LLMs do a great job because they understand both code and SVGs well. Edit: An example for a synth I'm buulding: https://imgur.com/a/U694Ek7 The irony of having to post it as a PNG isn't lost on me...

[deleted]

Re: Are AI labs pelicanmaxxing?

#166
post #154

Earlier quoted context omitted.

For example I do SVG diagrams to explain architecture and specific data flows in a multi-tier app. LLMs do a great job because they understand both code and SVGs well. Edit: An example for a synth I'm buulding: https://imgur.com/a/U694Ek7 The irony of having to post it as a PNG isn't lost on me...

This is not a strong use case. I could have made this in Microsoft word when I was 13 years old. Why on earth would you let an LLM do this when it would take you 10 min to do this in Figma or Inkscape or even just Word

Because it reads the code and generates this in a single pass in 5 minutes.

I haven't read the code.

Why would I do it in a slower, more difficult way for something that's going to be outdated in 2 hours?

Re: Are AI labs pelicanmaxxing?

#167

> All 21 pelican-bicycle images, across all seven labs, face right. No other animal/vehicle combination does that. > However, facing right is common: 60% of all 1,008 images do it. How common depends on the animal and the vehicle, and bicycles are one of the two vehicles where it’s strongest Of course the pelican on the bicycle is facing right. The drivetrain on a bicycle is on the right side. If you want any represe…

Interestingly, there was an artist a few years back who (for an unrelated project) had almost 400 people across a range of ages draw a bicycle and 75% of those faced left! So this seems to actually go slightly against the human drawing intuition. https://www.gianlucagimini.it/portfolio-item/velocipedia/ On the other hand, I notice that the prompt says to draw a pelican riding a bicycle, implying motion... and since m…

That's interesting. I think the bikes-facing-left bias probably comes from how humans use bikes: the kickstand is on the left, so people likely hold and approach bikes from its left.

Looking online, the kickstand is apparently on the left to avoid the gears. Based on the other comment about bike photography, it's interesting the same design choice makes humans and cameras/LLMs see bikes from different sides.

Re: Are AI labs pelicanmaxxing?

#168
post #59
post #44

Earlier quoted context omitted.

What if they’re not pelicanmaxxing, but svgmaxxxing in general? Because otherwise using a LLM to generate complex svgs is pretty niche and what I thought made this a good benchmark when it was new - generalized programming and spatial knowledge. Obviously image gen in svg format is not a particularly hard problem if tackled directly on its own.

But that's a genuine worthwhile capability. It's like benchmarkmaxxing on a weightlifting competition by getting really strong.

> benchmarkmaxxing on a weightlifting competition

If you stick to the benchpress, it's just "benchmaxxing".

Re: Are AI labs pelicanmaxxing?

#169
post #154

Earlier quoted context omitted.

For example I do SVG diagrams to explain architecture and specific data flows in a multi-tier app. LLMs do a great job because they understand both code and SVGs well. Edit: An example for a synth I'm buulding: https://imgur.com/a/U694Ek7 The irony of having to post it as a PNG isn't lost on me...

This is not a strong use case. I could have made this in Microsoft word when I was 13 years old. Why on earth would you let an LLM do this when it would take you 10 min to do this in Figma or Inkscape or even just Word

I'd be very impressed by anyone who can produce that in Figma or Word in ten minutes.

Re: Are AI labs pelicanmaxxing?

#170
post #169

Earlier quoted context omitted.

This is not a strong use case. I could have made this in Microsoft word when I was 13 years old. Why on earth would you let an LLM do this when it would take you 10 min to do this in Figma or Inkscape or even just Word

I'd be very impressed by anyone who can produce that in Figma or Word in ten minutes.

These are rectangles with some fill and strokes, textboxes lines and arrows. If the text has already been written, and the author has a clear idea of the flow all it takes are some a dozen or so points and click, and setting a couple of fills, then copy pastes, and finally some resizes, selects and moves.

This design is trivial. I admit it would be hard to achieve in Word (or at least for me because I don’t know how to make a diagram in Word any more) but Figma and Inkascape are made to do these things, and have optimized UI for that (personally I would have just used mermaid though).

I think you may have lost your faith in human capabilities just a little bit if you think drawing stuff like this takes any time or effort at all. Compared to designing and architecturing the system, drawing the diagram is trivial. Now I know that my parent did neither but it seems like they vibe-coded the whole thing. I’m sure they will end up with a fun little toy from the whole endeavor they can play with for 2 weeks before abandoning. Maybe the author will even feel bad about the carbon footprint of this whole exercise and buy some carbon offsets to make up for it.

Post reply on HN