Live data from Hacker News

Karpathy’s Pelican

twitter.com

251–260 of 461 posts

Re: Karpathy’s Pelican

#251

I don't think it's a bad way to benchmark new models, I just find it concerning that the author implies that "pelican on a bicycle" has been exhausted. At the risk of making overly broad, unfalsifiable claims I think multi-year exposure to AI content has dramatically raised our expectations for speed and volume but lowered them for quality. We see a very janky pelican and declare the problem solved.

Agreed. It's far from solved. Modern LLMs still generate pelican bike SVGs with obvious errors:

* some omitted the bottom of the diamond which connects from the pedals to the rear wheel

* some added an extra connection from the pedals to the front wheel, making it impossible to steer

* none could align the head tube with the fork

* none added a correct offset to the fork

* none could generate the chain properly in a way that attaches to the two sprockets correctly

I mean just look at these:

* Grok 4.5: https://s3.eu-west-1.amazonaws.com/images.dylancastillo.co/p...

* GPT 5.6 Terra: https://s3.eu-west-1.amazonaws.com/images.dylancastillo.co/p...

* Sonnet 5: https://s3.eu-west-1.amazonaws.com/images.dylancastillo.co/p...

from https://dylancastillo.co/posts/pelicanmaxxing.html

Re: Karpathy’s Pelican

#252

I can forgive the modeling being godawful jank (windows floating in the air, disconnected from the house). But I expected it to have a better understanding of the text. Instead, we have Bilbo's "disappearance" interpreted as him magically transporting or cloaking, and similarly for his reappearance.

But this only seems wrong to you because you're familiar with the prior context. When there's only a single paragraph to work from, and it's a drily humorous text, why not employ comic literalism and lean into the perplexity with which his neighbors viewed him?

The fact that a ring appears at the end proves that the model has context beyond what was in the prompt. The model clearly knows the story that the prompt was extracted from.

Re: Karpathy’s Pelican

#253

I don't think it's a bad way to benchmark new models, I just find it concerning that the author implies that "pelican on a bicycle" has been exhausted. At the risk of making overly broad, unfalsifiable claims I think multi-year exposure to AI content has dramatically raised our expectations for speed and volume but lowered them for quality. We see a very janky pelican and declare the problem solved.

Can someone explain what the pelican on a bicycle tests exactly? And why is it so important? I've never understood how it could translate to a useful task in real life.

Re: Karpathy’s Pelican

#254

> sure, why not, it's ~free Yeah, please check in with the folks protesting data center builds in their town causing their electricity prices to skyrocket and tap water to turn into a scarce resource. And now, instead of actually doing the above, please go ahead and downvote me, because how dare he question those LLM games.

> because how dare he question those LLM games. I thought you are the one who's on Greta's side, you ar e the one who should be saying "how dare you?"

Re: Karpathy’s Pelican

#255

I don't think it's a bad way to benchmark new models, I just find it concerning that the author implies that "pelican on a bicycle" has been exhausted. At the risk of making overly broad, unfalsifiable claims I think multi-year exposure to AI content has dramatically raised our expectations for speed and volume but lowered them for quality. We see a very janky pelican and declare the problem solved.

I think it's touching the limit of what one can reasonably expect any intelligent thing to produce with the only direction being "produce an svg of a pelican riding a bicycle". When you aren't sure if an LLM can write an svg well, or that it will be able to form a pelican shape, or animate a bicycle, it's a good test. After that, it's all judgement: how detailed should the pelican be? pelicans are the wrong shape for…

> touching the limit of what one can reasonably expect

If the expectation is that AI is going to replace "knowledge workers" then the limit would be a darn perfect drawing. We are nowhere close to that.

And Elon is already propagating the age of abundance where money won't exist anymore, right before calling the interviewing journalist dishonest and deservedly losing public trust. Smh my head.

Re: Karpathy’s Pelican

#256

I don't think it's a bad way to benchmark new models, I just find it concerning that the author implies that "pelican on a bicycle" has been exhausted. At the risk of making overly broad, unfalsifiable claims I think multi-year exposure to AI content has dramatically raised our expectations for speed and volume but lowered them for quality. We see a very janky pelican and declare the problem solved.

Can someone explain what the pelican on a bicycle tests exactly? And why is it so important? I've never understood how it could translate to a useful task in real life.

If your school mascot is a pelican and you need to make a flyer for the upcoming bike safety event, then this precise thing becomes useful. Most things that are useful in the real world seem not useful out of context.

Re: Karpathy’s Pelican

#257

I don't think it's a bad way to benchmark new models, I just find it concerning that the author implies that "pelican on a bicycle" has been exhausted. At the risk of making overly broad, unfalsifiable claims I think multi-year exposure to AI content has dramatically raised our expectations for speed and volume but lowered them for quality. We see a very janky pelican and declare the problem solved.

Humans are drawing pelicans riding bicycles now. Just google it and you will find 5 or 10 of them in the first few results. Including a t-shirt design. So it's a pretty much pointless test now.

Yet no LLM can actually do it.

It's quite surprising actually.

Re: Karpathy’s Pelican

#258

Earlier quoted context omitted.

Can someone explain what the pelican on a bicycle tests exactly? And why is it so important? I've never understood how it could translate to a useful task in real life.

If your school mascot is a pelican and you need to make a flyer for the upcoming bike safety event, then this precise thing becomes useful. Most things that are useful in the real world seem not useful out of context.

Yeah, pelican on a bicycle tells you how well the model can extrapolate outside of existing data, rather than just interpolate between it.

Re: Karpathy’s Pelican

#259

Earlier quoted context omitted.

Because the degree of realism in a video is determined by implicit knowledge of concepts like "in front of", "behind", "next to", "inside", "outside", "between", "occluded by", etc. as well as distances, angles, and the relative size of objects when viewed from a perspective.

I guess I just don't have an intuition for why that is different from what it does for non-"spatial" things. Like the fact that it works is still because it produced the code it did token by token. Whether its dealing with, e.g., " beside the rock" or " in the array", it's doing the same kind of inferential activity.

You are correct. It’s important not to let hype peddlers imply AI possesses consciousness.

Re: Karpathy’s Pelican

#260

This is just a very rough basic game skeleton. Companies believing AI is smart enough to spit out almost ready product will learn a hard way that it will will give them just a starting point that still needs a lot of work. This is a nature of all modern LLMs - they easily lose context: the bigger the context the more losses and distortions are, especially in the middle. So, giving AI the complete book does not mean i…

No, it cost 1M tokens, or about $10 EDIT: The parent comment originally claimed they spent $1M on the demo -- they seem to have edited it after I replied.

How many more millions of tokens do you need to spend to get something of acceptable quality?
Post reply on HN