Earlier quoted context omitted.
Aren't they still bad at understanding how bicycle frame works? Especially the steering part?
Maybe that makes them human.. https://www.booooooom.com/2016/05/09/bicycles-built-based-on...
Karpathy’s Pelican
331–340 of 461 posts
Re: Karpathy’s Pelican
#332Earlier quoted context omitted.
Maybe that makes them human.. https://www.booooooom.com/2016/05/09/bicycles-built-based-on...
It's funny to me, because I have a very strong spacial sense, and a lifetime of riding bikes. To the limit of my ability to draw, I can draw a flawless bicycle, down to the interlocked path the chain takes and the hanger on the derailleur.
Re: Karpathy’s Pelican
#333Earlier quoted context omitted.
Maybe that makes them human.. https://www.booooooom.com/2016/05/09/bicycles-built-based-on...
Humans aren't machines trained on the entire stolen corpus of human knowledge. We expect that a human will do poorly at arbitrary tasks they have no experience doing – especially drawing, which many (most?) humans aren't trained in at all. The same isn't true of AI, where its proponents, priests and proselytizers have spent time, energy and billions of dollars attempting to convince us it can do anything better than…
Elon Musk on Grok: “better than PhD level in everything.”
Re: Karpathy’s Pelican
#334I worked with an LLM to build a ~3D animation of the Back to the Future delorean Time Machine as a way to spice up the hero on a docs page. That took a fair amount of custom tuning and I had to create a tuning view to get some of the behaviors right. But it was enough fun that I generalized it to take in ~any scene description from a film. It goes out and gets more detailed descriptions and film stills if available b…
I am interested, I'd love to see! I'm waiting for the day my dad's self published books become self produced movies!
Re: Karpathy’s Pelican
#335Opus 5 is the first I’ve seen to “one shot” it (in a harness, so it was more than one LLM call).
Re: Karpathy’s Pelican
#336It's difficult to think in exponentials. But this demonstrates we're a couple orders of magnitude away from generating 1:1 hyper-personalized entertainment and media for individuals, rather than the masses.
I don't really get the desire for hyper-personalized entertainment. People are very good at pointing out things they dislike, but not very good at coming up with how to fix them (common wisdom in game design). On top of that, a decent chunk of the joy of entertainment is the social aspect.
Now, whether this would be at all possible, or even healthy for you, I don't know.
Re: Karpathy’s Pelican
#337Re: Karpathy’s Pelican
#338Earlier quoted context omitted.
Agree. The pelican benchmark was interesting a year ago when most models struggled and a good pelican indicated an unusually capable model. Now it’s saturated and uninteresting. A good new benchmark should have awful performance to start and there should be a lot of headroom for improvement. This benchmark is also intentionally difficult and requires the LLM to develop the animation through spatial reasoning and firs…
I think it's interesting to see them visibly struggling to improve. Claude pelicans aren't much better today than they where 18 months.
Re: Karpathy’s Pelican
#339Last I checked, I did not see the prompt for this really cool thing, so it is not reproducible.
Did I miss the prompt somewhere?
Re: Karpathy’s Pelican
#340Earlier quoted context omitted.
> If the computer can't do it better than a human being, then what's the point? Because the benchmark wasn't testing "can an LLM draw a pelican like a human". The original article was testing the relative capabilities between LLMs. Now that LLMs can all draw pelicans all similarly, the test is less interesting as a comparative benchmark.
Trillions of dollars spent. Trillions of gigawatts consumed. And people still celebrate "Yay! We're less wrong than the other guys!" This is what the tech industry has become? Less of a failure is still failure.
LLM has progressed a lot in the last two year, judging from the pelican drawings. I personally couldn't care less about it though. I do know that I've gone from using no AI at all for coding to probably 95%. I hardly code by hand anymore. That's much more impressive and significant. Failure you said?