I still believe my bench is superior to Pelican or Karpathy’s. Everyone knows how a MacBook looks. Any missing or incorrect detail will be obvious. However, Pelican or this world can be anything. https://playcode.io/blog/macbook-svg-benchmark
Karpathy’s Pelican
271–280 of 461 posts
Re: Karpathy’s Pelican
#272IMO, the area where AI is going to be most useful over the next couple years is in developing manufacturing processes top to bottom. Maybe a million token budget is too small, but something like "design me a sneaker and all the equipment to manufacture it autonomously".
Gotta love when a techbro just says some complete nonsense like this with total confidence
Re: Karpathy’s Pelican
#273I don't think it's a bad way to benchmark new models, I just find it concerning that the author implies that "pelican on a bicycle" has been exhausted. At the risk of making overly broad, unfalsifiable claims I think multi-year exposure to AI content has dramatically raised our expectations for speed and volume but lowered them for quality. We see a very janky pelican and declare the problem solved.
Can someone explain what the pelican on a bicycle tests exactly? And why is it so important? I've never understood how it could translate to a useful task in real life.
It used to be a very difficult task for models, see [2,3,4]
it cuts across several tasks that AI used to be very bad at, but now has improved quite a bit. Namely, spatial reasoning (because it has to manually place the points of the svg such that they make sense and form what it says it forms. This used to not work very well, with random shapes floating around that it would mark things like "eyebrows" but were nowhere near the "eyes", etc.
It also tests the model's world knowledge (what do pelicans look like? sure they have wings, feet, beaks etc, but what shape are they? how to get proportions roughly right? this isn't a given from text data about the bird. This goes doubly for a bike, which is a quite complex shape that most humans fail to draw correctly[1] (many draw the frame or chain connecting in impossible ways that would not ever function mechanically)
Before it was pelican on a bicycle there were people having it do horses/unicorns making the rounds - gpt4.0 or whatever would often make hideous abominations of legs and mouths
[1] https://www.gianlucagimini.it/portfolio-item/velocipedia/
[2] https://static.simonwillison.net/static/2026/mistral-small-4...
[3] https://static.simonwillison.net/static/2025/codex-hacking-m...
[4] https://static.simonwillison.net/static/2025/gemini-2.5-flas...
Re: Karpathy’s Pelican
#274Earlier quoted context omitted.
I'm mostly seeing the impact of GenAI images in everydays life for example in the small posters small associations or individuals usually attach on streets to promote some small local event. We went from just text created with PowerPoint and maybe some stock image just a Google search away to now images depicting the topic closely. But yeah, it's not like a revolution. And this personal media thing, yeah maybe for te…
I yearn for the old days of bad photoshop/wordart/powerpoint flyers. They were visually bad but honest and sometimes soulful or playful. The AI slop that's everywhere now looks superficially more professional but it's very busy, samey, unnatural and it really turns me off.
Re: Karpathy’s Pelican
#275Earlier quoted context omitted.
Better than the mean human? The mean human creates far better stories while daydreaming. AI enthusiasts have such a distorted view of human capability, it’s bizarre.
The mean human gets a middling score in creative writing tasks at the end of their mandatory education, and most then leave school and forget what little they ever learned outside whatever their career path happened to be. Most humans never do a creative writing course after school, and the longest fiction most people will write is their resume description of what their previous jobs involved, or perhaps their dating…
Re: Karpathy’s Pelican
#276I'd rather have them battle on the topic "Who builds a better Google Wave for LLM chats" to explore the space of how AI studios could be. "Getting Started with Google Wave": https://www.youtube.com/watch?v=eKUAqNGVwX0
https://www.youtube.com/watch?v=4Z4RKRLaSug
Their "World Quester 2" tutorial shines as the holy grail of consistent and ergonomic user interface and game design. The menuing system is so magnificently structured and well organized, it bring tears to my eyes. Google Wave pales in comparison.
Re: Karpathy’s Pelican
#277Earlier quoted context omitted.
No, it cost 1M tokens, or about $10 EDIT: The parent comment originally claimed they spent $1M on the demo -- they seem to have edited it after I replied.
How many more millions of tokens do you need to spend to get something of acceptable quality?
Re: Karpathy’s Pelican
#278Earlier quoted context omitted.
I guess I just don't have an intuition for why that is different from what it does for non-"spatial" things. Like the fact that it works is still because it produced the code it did token by token. Whether its dealing with, e.g., " beside the rock" or " in the array", it's doing the same kind of inferential activity.
You are correct. It’s important not to let hype peddlers imply AI possesses consciousness.
Re: Karpathy’s Pelican
#279Earlier quoted context omitted.
I haven't really seen evidence that any ai can reliably draw a pelican riding a bicycle. Not if you look at the image long enough to take it in. Even the best ones have something wrong with them. Not a matter of taste but a matter of having both legs peddling on the viewer's side of the bicycle or having two beaks. I'm actually beginning to wonder if some people who ignore these things have a different, somewhat less…
> I haven't really seen evidence that any ai can reliably draw a pelican riding a bicycle Please remember, we've started from there : https://simonwillison.net/2024/Oct/25/pelicans-on-a-bicycle/ When it started, it was clear what LLM would stand out, its style, etc. Nowadays, the pelicans look similar, the difference is in details and sometimes hard to catch. Sure, the task is not completed perfectly, but that's not…
Isn't it?
If the computer can't do it better than a human being, then what's the point?
Being wrong at scale is not better than being right.
Re: Karpathy’s Pelican
#280That method of perception probably scales N^2... so sure with more compute, LoTR animation will improve. But I think to get a real jump in "experiential feedback", perception needs to scale linear or sublinear. Maybe that's there LeCunn's jepa will come in.
There needs to be the removal of the middle man:
image -> text -> action
To image -> action.