Live data from Hacker News

Karpathy’s Pelican

twitter.com

331–340 of 461 posts

Re: Karpathy’s Pelican

#331
post #165
post #158

Earlier quoted context omitted.

Aren't they still bad at understanding how bicycle frame works? Especially the steering part?

Maybe that makes them human.. https://www.booooooom.com/2016/05/09/bicycles-built-based-on...

It's funny to me, because I have a very strong spacial sense, and a lifetime of riding bikes. To the limit of my ability to draw, I can draw a flawless bicycle, down to the interlocked path the chain takes and the hanger on the derailleur.

Re: Karpathy’s Pelican

#332
post #165

Earlier quoted context omitted.

Maybe that makes them human.. https://www.booooooom.com/2016/05/09/bicycles-built-based-on...

It's funny to me, because I have a very strong spacial sense, and a lifetime of riding bikes. To the limit of my ability to draw, I can draw a flawless bicycle, down to the interlocked path the chain takes and the hanger on the derailleur.

I ride a recumbent trike, and the chain path is even more grotesque. I think I could draw it.

Re: Karpathy’s Pelican

#333
post #165

Earlier quoted context omitted.

Maybe that makes them human.. https://www.booooooom.com/2016/05/09/bicycles-built-based-on...

Humans aren't machines trained on the entire stolen corpus of human knowledge. We expect that a human will do poorly at arbitrary tasks they have no experience doing – especially drawing, which many (most?) humans aren't trained in at all. The same isn't true of AI, where its proponents, priests and proselytizers have spent time, energy and billions of dollars attempting to convince us it can do anything better than…

Sam Altman: “GPT-5 is the first time that it really feels like talking to an expert in any topic, like a PhD-level expert.”

Elon Musk on Grok: “better than PhD level in everything.”

Re: Karpathy’s Pelican

#334
post #8

I worked with an LLM to build a ~3D animation of the Back to the Future delorean Time Machine as a way to spice up the hero on a docs page. That took a fair amount of custom tuning and I had to create a tuning view to get some of the behaviors right. But it was enough fun that I generalized it to take in ~any scene description from a film. It goes out and gets more detailed descriptions and film stills if available b…

I am interested, I'd love to see! I'm waiting for the day my dad's self published books become self produced movies!

Depending on how deep you want to go down the rabbit hole, that's possible-ish now: https://www.youtube.com/watch?v=V5xi3Usi0i8

Re: Karpathy’s Pelican

#335
A simple prompt that still stumps frontier LLMs most of the time is “create a pinball game”. They’ll put all the right pieces there but then fail to arrange them such that the game is truly playable. They’ll put a wall in the way of the launch chute so the ball can’t be launched. Or the flippers will pivot the wrong way. Or there will be holes such that the ball drops off the bottom without getting within reach of the flippers, etc.

Opus 5 is the first I’ve seen to “one shot” it (in a harness, so it was more than one LLM call).

Re: Karpathy’s Pelican

#336
post #73

It's difficult to think in exponentials. But this demonstrates we're a couple orders of magnitude away from generating 1:1 hyper-personalized entertainment and media for individuals, rather than the masses.

I don't really get the desire for hyper-personalized entertainment. People are very good at pointing out things they dislike, but not very good at coming up with how to fix them (common wisdom in game design). On top of that, a decent chunk of the joy of entertainment is the social aspect.

The only charitable interpretation of the hyper-personalization take I can give, is that you would want to see content that elicited a certain type of emotion inside you (so, a very high-level description).

Now, whether this would be at all possible, or even healthy for you, I don't know.

Re: Karpathy’s Pelican

#338

Earlier quoted context omitted.

Agree. The pelican benchmark was interesting a year ago when most models struggled and a good pelican indicated an unusually capable model. Now it’s saturated and uninteresting. A good new benchmark should have awful performance to start and there should be a lot of headroom for improvement. This benchmark is also intentionally difficult and requires the LLM to develop the animation through spatial reasoning and firs…

I think it's interesting to see them visibly struggling to improve. Claude pelicans aren't much better today than they where 18 months.

The code I see is a lot like the pelicans. All of the code in codebases, good, bad and ugly, is slowly being replaced by whatever level of code ai is currently able to create. All code is now a slightly wonky pelican on a bike, but if you look closely, it doesn’t fully make sense. Since ai is converging on less wonky, but not internally consistent, we’re just moving on to what is possible with high volume instead of detailed quality. I think that is the ai software world as well.

Re: Karpathy’s Pelican

#339
The difference between this and Simon's pelican is that with Simon, I get the prompt.

Last I checked, I did not see the prompt for this really cool thing, so it is not reproducible.

Did I miss the prompt somewhere?

Re: Karpathy’s Pelican

#340

Earlier quoted context omitted.

> If the computer can't do it better than a human being, then what's the point? Because the benchmark wasn't testing "can an LLM draw a pelican like a human". The original article was testing the relative capabilities between LLMs. Now that LLMs can all draw pelicans all similarly, the test is less interesting as a comparative benchmark.

Trillions of dollars spent. Trillions of gigawatts consumed. And people still celebrate "Yay! We're less wrong than the other guys!" This is what the tech industry has become? Less of a failure is still failure.

I don't know how much money has been spent for AI, and I very much doubt you do. Do you know if more has been spent on LLM the last 9 years -- since "Attention is all you need" -- than Internet infrastructure during, say 1995 to 2004? That included the dotcom crash. Did you lament how a failure the Internet was?

LLM has progressed a lot in the last two year, judging from the pelican drawings. I personally couldn't care less about it though. I do know that I've gone from using no AI at all for coding to probably 95%. I hardly code by hand anymore. That's much more impressive and significant. Failure you said?

Post reply on HN