Live data from Hacker News

Karpathy’s Pelican

twitter.com

411–420 of 461 posts

Re: Karpathy’s Pelican

#411
post #327

Earlier quoted context omitted.

What does this even test for? Can I use LLMs to directly generate machine code to replace my compiler? Or maybe I can use LLMs as a bare metal OS / scheduler to replace my machine's operating system and scheduler? It makes zero sense to test for that. Not only that this so-called "benchmark" isn't economically useful, but that it tests for the sake of testing; and for attention. To end this obsession with generating…

It tests for the "I" in "AI".

So if we test models that only output text to directly generate waveforms of sound from text or binary code, or directly generating binary code to replace a compiler, does that mean it is "intelligent"?

Does that even test for intelligence?

This is like testing if a horse can fly just because someone showed an image of a Pegasus, or testing if a fish can climb up a tree and believing they are not intelligent because each of them cannot fly or climb up trees.

Re: Karpathy’s Pelican

#412

Earlier quoted context omitted.

he said the prompt was the first paragraph of LoTR, but he didn't mention a preamble this guy seems to have taken that idea and got something similar/better, so likely the prompt isn't too special https://x.com/Izkimar/status/2083819741643178208?s=20

That was far worse as far as visual story-telling. Anthropic wins, which is to be expected against whatever that product is. In either case, I still don't see how I could reproduce this to test against various models, which is the entire point of Simon's pelican.

The second one is also Anthropic, same model even.

Re: Karpathy’s Pelican

#413

A simple prompt that still stumps frontier LLMs most of the time is “create a pinball game”. They’ll put all the right pieces there but then fail to arrange them such that the game is truly playable. They’ll put a wall in the way of the launch chute so the ball can’t be launched. Or the flippers will pivot the wrong way. Or there will be holes such that the ball drops off the bottom without getting within reach of th…

We don't use these models by way of "one shot" so I don't see why it's relevant. It's clearly useful to let them iterate.

Re: Karpathy’s Pelican

#414

Earlier quoted context omitted.

Because the degree of realism in a video is determined by implicit knowledge of concepts like "in front of", "behind", "next to", "inside", "outside", "between", "occluded by", etc. as well as distances, angles, and the relative size of objects when viewed from a perspective.

I guess I just don't have an intuition for why that is different from what it does for non-"spatial" things. Like the fact that it works is still because it produced the code it did token by token. Whether its dealing with, e.g., " beside the rock" or " in the array", it's doing the same kind of inferential activity.

If you're at (-2, 3, 5), pointing in the (2, 1, -1) direction (as a direction vector), then which is in front of the other from your point of view: the object at (4, 5, -8) or the one at (6, 4, 3)? (or neither)

You can't just defer this decision to the three.js program because you need to understand this sort of relationship yourself if you're going to suitably place objects (and your own view) in a three.js scene in the first place. Building a complex scene involves making hundreds, or more, of decisions about where to place coordinates in 3d space.

This example is underspecified (e.g. what are the shapes and sizes of the objects and what is the angular width of your view) but it illustrates the problem. Even with a good intuitive understanding of 3d space, we would struggle with this. LLMs are not calculators, so it's surprising if they manage it.

Re: Karpathy’s Pelican

#415

Earlier quoted context omitted.

Pelicans don't ride bicycles. It's physically impossible. The problem is to draw it in the least disturbing way possible.

True! But somehow Disney has been drawing ducks riding bikes in a way that seems to satisfy everyone since before my grandfather was born. https://ridesabike.com/donald-duck-daisy-duck-huey-dewey-and...

It’s better than many of the AI offerings and the bike could steer and the ducks are sitting on saddles, but the three nephews can’t reach the bottom of the pedal stroke and by the looks of their feet on the far side of the bike they aren’t trying. That shouldn’t satisfy Donald and Daisy, leaving them with all the work.

Re: Karpathy’s Pelican

#416

Earlier quoted context omitted.

To help explain AI to my elderly mother, I said it is like an alien intelligence on a planet too far away to directly observe Earth, and everything it knows about humanity and our world it learned from reading just about every book and website. You can ask it a question or to do some work and it can often give a useful response but it can never directly judge its accuracy if it relates to the physical world, it has o…

That's a nice metaphor for your mother, i like it. I would just touch on one small part, which is that it has more than our judgement though. It has access to laws of physics, to formulas, to all our current scientific knowledge which we have shown to be correct by interacting with the physical world. It can create an (incomplete) model of the world based on verified theories already. Through deduction and reasoning…

That's true, on some level it is capable of reasoning about the physical world even beyond our ability, through sheer breadth of knowledge and stamina. I suppose the area where it relies most on our judgment is about subjective experience, and this better explains its limitations.

Re: Karpathy’s Pelican

#418

Earlier quoted context omitted.

It's funny to me, because I have a very strong spacial sense, and a lifetime of riding bikes. To the limit of my ability to draw, I can draw a flawless bicycle, down to the interlocked path the chain takes and the hanger on the derailleur.

I ride a recumbent trike, and the chain path is even more grotesque. I think I could draw it.

I also ride a recumbent: Catrike Expedition. I'm not sure I would get the subtleties of the chain guides correct. What's yours?

https://www.catrike.com/expedition

Re: Karpathy’s Pelican

#419

Earlier quoted context omitted.

he said the prompt was the first paragraph of LoTR, but he didn't mention a preamble this guy seems to have taken that idea and got something similar/better, so likely the prompt isn't too special https://x.com/Izkimar/status/2083819741643178208?s=20

That was far worse as far as visual story-telling. Anthropic wins, which is to be expected against whatever that product is. In either case, I still don't see how I could reproduce this to test against various models, which is the entire point of Simon's pelican.

you are seeing single realizations of non-deterministic processes. Confidently stating one model is better than the other is not really possible this way; that's just like stating one dice is better than the other because you rolled a six with that one on the first try.

Re: Karpathy’s Pelican

#420
post #93

A lot of people are posting here about how bad the end product is, but that is kind of the point. Models have moved beyond generating images to a new kind of benchmark that better exposes understanding of the physical world, and we can use benchmarks like this to measure future progress. (Of course, it will have to be a qualitative/subjective measurement.)

Will Smith spaghetti was garbage a couple of years ago and now AI videos are becoming close to indistinguishable from real videos in many cases. Spaghetti 2026: https://x.com/dreamingtulpa/status/2083304533829066873 https://xcancel.com/dreamingtulpa/status/2083304533829066873

I wonder though if models are now "benchmaxxing" against these kinds of prompts. I haven't needed to use them so I can't say, but would be interesting.
Post reply on HN