Live data from Hacker News

Karpathy’s Pelican

twitter.com

281–290 of 461 posts

Re: Karpathy’s Pelican

#281
post #159
post #3

I'd rather have them battle on the topic "Who builds a better Google Wave for LLM chats" to explore the space of how AI studios could be. "Getting Started with Google Wave": https://www.youtube.com/watch?v=eKUAqNGVwX0

It’s a fun idea to re-animate dead google products by feeding product videos to an AI

Remember when Robert Scoble posted a photo of himself naked in the shower wearing only Google Wave?

Oh never mind, that was Google Glass, another dead google product. That he personally killed with that photo. So confusing to keep track of them all.

Re: Karpathy’s Pelican

#283

I don't think it's a bad way to benchmark new models, I just find it concerning that the author implies that "pelican on a bicycle" has been exhausted. At the risk of making overly broad, unfalsifiable claims I think multi-year exposure to AI content has dramatically raised our expectations for speed and volume but lowered them for quality. We see a very janky pelican and declare the problem solved.

I think it's touching the limit of what one can reasonably expect any intelligent thing to produce with the only direction being "produce an svg of a pelican riding a bicycle". When you aren't sure if an LLM can write an svg well, or that it will be able to form a pelican shape, or animate a bicycle, it's a good test. After that, it's all judgement: how detailed should the pelican be? pelicans are the wrong shape for…

> I think it's touching the limit of what one can reasonably expect any intelligent thing to produce with the only direction being "produce an svg of a pelican riding a bicycle".

Oh come now. I am extremely confident that if I hired a professional artist to draw a picture of a pelican riding a bicycle, I would get something inarguably much better than what today's best coding LLMs can produce.

Re: Karpathy’s Pelican

#284
How long we will be testing and benchmarking generative capabilities of LLMs? in creativity, in code generation less or more result is expected and approved, but in execution $19.8 can not be $19.9 or $19.7.

Re: Karpathy’s Pelican

#285
post #8

I worked with an LLM to build a ~3D animation of the Back to the Future delorean Time Machine as a way to spice up the hero on a docs page. That took a fair amount of custom tuning and I had to create a tuning view to get some of the behaviors right. But it was enough fun that I generalized it to take in ~any scene description from a film. It goes out and gets more detailed descriptions and film stills if available b…

I am interested, I'd love to see! I'm waiting for the day my dad's self published books become self produced movies!

Thank you, for your patience in my reply here it is:

https://banagale.com/cinematic-canvas-ai-film-animation.htm

I provide a "how I got to this" up front, but if you want to jump right to the Apocalypto stuff, use this:

https://banagale.com/cinematic-canvas-ai-film-animation.htm#...

And if you want to play with interactive demos of the two animation sequences (the tuning tooling I described in my OP) you can go directly there:

https://banagale.com/cinematic-canvas-workbench-demos

If anyone wants to collaborate on building out the cinematic-canvas-workbench project please email me.

Re: Karpathy’s Pelican

#286

Earlier quoted context omitted.

I think it's touching the limit of what one can reasonably expect any intelligent thing to produce with the only direction being "produce an svg of a pelican riding a bicycle". When you aren't sure if an LLM can write an svg well, or that it will be able to form a pelican shape, or animate a bicycle, it's a good test. After that, it's all judgement: how detailed should the pelican be? pelicans are the wrong shape for…

> I think it's touching the limit of what one can reasonably expect any intelligent thing to produce with the only direction being "produce an svg of a pelican riding a bicycle". Oh come now. I am extremely confident that if I hired a professional artist to draw a picture of a pelican riding a bicycle, I would get something inarguably much better than what today's best coding LLMs can produce.

At that rate you could use an image model, which was designed for the task. Thats the absurdity of this test. It’s often a text only generation model that has never seen a pelican, coerced into creating an xml graphical representation that a human might recognize. That it does anything passable is already astounding.

Re: Karpathy’s Pelican

#287
post #270

Earlier quoted context omitted.

> I haven't really seen evidence that any ai can reliably draw a pelican riding a bicycle Please remember, we've started from there : https://simonwillison.net/2024/Oct/25/pelicans-on-a-bicycle/ When it started, it was clear what LLM would stand out, its style, etc. Nowadays, the pelicans look similar, the difference is in details and sometimes hard to catch. Sure, the task is not completed perfectly, but that's not…

Sure, the task is not completed perfectly, but that's not the point. Isn't it? If the computer can't do it better than a human being, then what's the point? Being wrong at scale is not better than being right.

Many humans would struggle with this even with very good tooling (ie not writing raw svg and using illustrator). I struggle to draw a bicycle accurately. But yes, I suspect it will be diminishing returns and I doubt it will ever be perfect due to the average nature of AI but I’d like to be wrong.

Re: Karpathy’s Pelican

#288

Earlier quoted context omitted.

Agree. The pelican benchmark was interesting a year ago when most models struggled and a good pelican indicated an unusually capable model. Now it’s saturated and uninteresting. A good new benchmark should have awful performance to start and there should be a lot of headroom for improvement. This benchmark is also intentionally difficult and requires the LLM to develop the animation through spatial reasoning and firs…

I think it's interesting to see them visibly struggling to improve. Claude pelicans aren't much better today than they where 18 months.

Rendering 3d worlds has hugely improved though.

Re: Karpathy’s Pelican

#289
post #270

Earlier quoted context omitted.

> I haven't really seen evidence that any ai can reliably draw a pelican riding a bicycle Please remember, we've started from there : https://simonwillison.net/2024/Oct/25/pelicans-on-a-bicycle/ When it started, it was clear what LLM would stand out, its style, etc. Nowadays, the pelicans look similar, the difference is in details and sometimes hard to catch. Sure, the task is not completed perfectly, but that's not…

Sure, the task is not completed perfectly, but that's not the point. Isn't it? If the computer can't do it better than a human being, then what's the point? Being wrong at scale is not better than being right.

>If the computer can't do it better than a human being, then what's the point?

Because the benchmark wasn't testing "can an LLM draw a pelican like a human". The original article was testing the relative capabilities between LLMs. Now that LLMs can all draw pelicans all similarly, the test is less interesting as a comparative benchmark.

Re: Karpathy’s Pelican

#290
post #270

Earlier quoted context omitted.

I haven't really seen evidence that any ai can reliably draw a pelican riding a bicycle. Not if you look at the image long enough to take it in. Even the best ones have something wrong with them. Not a matter of taste but a matter of having both legs peddling on the viewer's side of the bicycle or having two beaks. I'm actually beginning to wonder if some people who ignore these things have a different, somewhat less…

> I haven't really seen evidence that any ai can reliably draw a pelican riding a bicycle Please remember, we've started from there : https://simonwillison.net/2024/Oct/25/pelicans-on-a-bicycle/ When it started, it was clear what LLM would stand out, its style, etc. Nowadays, the pelicans look similar, the difference is in details and sometimes hard to catch. Sure, the task is not completed perfectly, but that's not…

When is it ever hard to catch?
Post reply on HN