Live data from Hacker News

Are AI labs pelicanmaxxing?

dylancastillo.co

141–150 of 253 posts

Re: Are AI labs pelicanmaxxing?

#141

Earlier quoted context omitted.

I don't think this small amount generalization to other animals and vehicles is strong evidence they haven't trained on this, either directly or more generally.

Honest question how could they possibly train on this as there are no good SVG pelicans to train off right? So they’re just training off a bunch of bad ones which should lead to just bad pelicans, but the pelicans are getting better.

Training on generating SVGs directly at all is already fairly niche. Generating full scenes with a cartoony character is even nicher. But there's plenty of non-pelican cartoony SVG content out there (created, not written, by humans with vector design tools), and more importantly, plenty of vision models to give feedback on the output (just raster as a png). You could easily hill climb this niche skill, if you cared.

Re: Are AI labs pelicanmaxxing?

#142
post #94
post #46

https://en.wikipedia.org/wiki/Goodhart%27s_law "Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes." Or the more pop layman version "When a measure becomes a metric/KPI, it ceases to be a good measure." Story time, I live in Argentina, and we don't have Big Macs, the main Mc Donald's brand, here, because during the CFK presidency, one of her tactics was to G…

How did the president manage to influence McDonalds' local business decisions? And how did that lead to McDonald's pulling out of the country?

McDonald's complies with local laws and operates with local juristic entities. While McDonald's US owns McDonalds Argentina (Arcos Dorados Argentina SA), and there's some level of control that occurs in the US "Global" headquarters, some that occurs at the National HQ level, and some that occurs in the individual location. Obviously a president would be able to influence 2 of those levels, it's not necessary for her to influence the Global executives.

To be clear McDonald's didn't pull out of the country, like in many other countries, it adapted to the local market. Same way as in india it serves non meat variants owing to the high vegetarian population, consider that mcdonald's operates in Venezuela and China, and operated in Russia up until the Russia-Ukraine war. It takes a lot for MCD to pull out of a country.

As for the exact mechanisms of presidential price control on MCD Argentina, I don't have the specifics here, but I can get pretty close. There have been 2 broad mechanisms to exhert price control during CFK's 8 years of presidency and during his Husband's 4, let's call them official and unofficial.

The official mechanisms would be passing laws or presidential decrees (DNU) that don't go through congress, as well as influencing regulation of executive ministries/departments like central bank norms, exchange rates. Some measures like 'precios cuidados' were placed for this very specific purpose, it's very possible Big Macs were under this specific scheme, I do not recall, I was a bit young.

The unofficial methods would be less public, but well known, in the food industry it was especially common, it's well established that food are one of the first and most common targets of price controls. I have heard direct accounts from family members about high ranking government officials setting up meetings with producers in food markets to give orders of lowering prices, one going as far as brandishing a firearm by placing it on a table while discussing the subject.

Writing this out loud I realize that this explains the later over-correction of argentina that allowed freemarket capitalist libertarianism to rise. The optimal strategy in democracy seems to be polarization, so both extremisms seem to symbiotically feed off each other. Not a country of moderateness this one.

Re: Are AI labs pelicanmaxxing?

#143

> All 21 pelican-bicycle images, across all seven labs, face right. No other animal/vehicle combination does that. > However, facing right is common: 60% of all 1,008 images do it. How common depends on the animal and the vehicle, and bicycles are one of the two vehicles where it’s strongest Of course the pelican on the bicycle is facing right. The drivetrain on a bicycle is on the right side. If you want any represe…

Interestingly, there was an artist a few years back who (for an unrelated project) had almost 400 people across a range of ages draw a bicycle and 75% of those faced left! So this seems to actually go slightly against the human drawing intuition. https://www.gianlucagimini.it/portfolio-item/velocipedia/ On the other hand, I notice that the prompt says to draw a pelican riding a bicycle, implying motion... and since m…

It’s a very strong convention in bicycle photography. The drive side is considered nicer looking, it’s where you get to show off the components, etc.

Independent of what humans may do with the same prompt, training datasets will be full of traditional bike photos, which will all be drive side.

It’s also one of the easiest yellow flags to look for in used bike listings (not low-end, but anything enthusiast level). If a bike is photographed on the wrong side, there’s a non-negligible chance it’s stolen, because the seller is presumed to know better if they’re an enthusiast themself.

Re: Are AI labs pelicanmaxxing?

#144
Bike nerd + AI nerd here. The author's observation that all bicycle images face right almost certainly has to do with the convention to photograph a bicycle from the right. From the right, you see the drivetrain - this is good for aesthetics, but also for marketing - the drivetrain is branded and labelled and a buyer will want to know what model it is.

There is a bunch of guidance online on how to photograph bikes, and every sales image of a bike will be from the right. You can anecdotally observe this by google imaging 'bicycle for sale'.

Re: Are AI labs pelicanmaxxing?

#145
post #17

This is fantastic I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations. Catching a lab cheating specifically on my one dumb benchmark would be really funny . Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than a…

I worked with image analytics years ago. In 2024 and cynical about LLM image capability, I tested different models to create "Two female octopi in a bar each having their own drink. One is wearing a yellow hat. The other is wearing a red hat. They have no human features." It went badly I think because the training visual training data associated with the word "female" was overwhelmingly biased in quantity and obviated the 'no human features' instruction even after iterated instructions to remove.

So I tried "Make an image of a Djibouti cab with a camel sitting in the passenger seat. Give the camel no human anatomical features." Still bad.

Two years later the graphics and perspectives of the tools are so much better. But that isn't the big story. What really stands out is the models are no longer adding human female anatomies to camels and octopi. The gates and filtering based on instructions have matured enough so that the pelican/bicycle deductive reasoning puzzle is less problematic. But until an LLM anticipates something like impressionism from Parisian artists rebelling against the rules of the French Academy, continue to reserve a place for human artistry.

tl;dr simonw, your spot-check nailed it just as well as any extensive methodology. LLM's no longer need to cheat this part of the test (better to hack the question than the tool).

Re: Are AI labs pelicanmaxxing?

#146

> All 21 pelican-bicycle images, across all seven labs, face right. No other animal/vehicle combination does that. > However, facing right is common: 60% of all 1,008 images do it. How common depends on the animal and the vehicle, and bicycles are one of the two vehicles where it’s strongest Of course the pelican on the bicycle is facing right. The drivetrain on a bicycle is on the right side. If you want any represe…

Interestingly, there was an artist a few years back who (for an unrelated project) had almost 400 people across a range of ages draw a bicycle and 75% of those faced left! So this seems to actually go slightly against the human drawing intuition. https://www.gianlucagimini.it/portfolio-item/velocipedia/ On the other hand, I notice that the prompt says to draw a pelican riding a bicycle, implying motion... and since m…

I wonder how the skill at drawing a bicycle compares between the general population, and those who own or ride one regularly.

Re: Are AI labs pelicanmaxxing?

#147
post #59

Earlier quoted context omitted.

But that's a genuine worthwhile capability. It's like benchmarkmaxxing on a weightlifting competition by getting really strong.

How useful actually is this? It generates SVGs of pelicans on bicycles, sure, and some of them are (almost) spatially correct. But, none of them look good . AI image generation suffers from this more generally. You can generate pictures of pelicans, sure. Newer models clearly generate images with more pelican-ness than before. But all of it is still uglier than sin. Drawing things accurately is one thing, making resu…

Would it be useful if the models were actually good at it?

Re: Are AI labs pelicanmaxxing?

#148

Earlier quoted context omitted.

I don't think this small amount generalization to other animals and vehicles is strong evidence they haven't trained on this, either directly or more generally.

Honest question how could they possibly train on this as there are no good SVG pelicans to train off right? So they’re just training off a bunch of bad ones which should lead to just bad pelicans, but the pelicans are getting better.

They can easily afford 1000 human made pelican svg files if they want. I think you underestimate how much resources SotA AI companies have.

(I'm not saying that they did that. I'm just saying they can.)

Re: Are AI labs pelicanmaxxing?

#149
post #59

Earlier quoted context omitted.

But that's a genuine worthwhile capability. It's like benchmarkmaxxing on a weightlifting competition by getting really strong.

The fear is that the SVGmaxxing is limited to "X doing Y". If such 'template maxxing' exists, it will break for other templates, e.g. "X not doing Y", "X and Y doing Z", "X doing Y doing Z", etc.

If a model can improve at drawing "X doing Y" and that prompt wasn't in the training set then it means it has improved its internal mapping from text-to-spatial-to-text.
Post reply on HN