Live data from Hacker News

Karpathy’s Pelican

twitter.com

431–440 of 461 posts

Re: Karpathy’s Pelican

#431
post #262

Earlier quoted context omitted.

I'd like to see tests of things the current AIs are bad at, like drive a car. Or maybe take instruction to complete some novel activity to test how well they can learn.

Would you say that the generated video is good? Better than I would have expected? Sure. Impressive that it got that far? Definitely. But is the end result good?

The results are fine, it just isn't a very interesting benchmark. It can generate video and it's getting better. Great, I guess.

Clocking improvements in areas the AIs are terrible at (assuming the goal is still AGI) are more interesting.

Re: Karpathy’s Pelican

#434
I am amazed by Karpathy's writing. He can express complicated thoughts on these topics in a way that makes you feel you would have come to the same conclusion :)

Re: Karpathy’s Pelican

#435

Earlier quoted context omitted.

I guess I just don't have an intuition for why that is different from what it does for non-"spatial" things. Like the fact that it works is still because it produced the code it did token by token. Whether its dealing with, e.g., " beside the rock" or " in the array", it's doing the same kind of inferential activity.

That’s becoming as illuminating as saying you produced that comment word by word.

What is becoming illuminating?

Re: Karpathy’s Pelican

#436

Earlier quoted context omitted.

[flagged]

you're doing the exact same thing by making the inverse claim without any information, at least the original comment claims some change for someone to investigate... all you're doing is disagreeing as an opportunity to be snarky

Pseudo-intellectual drivel.

The OP is calling Karpathy out for producing slop, all because he loves playing with AI. Karpathy has done more for the industry than 10 of the OPs combined. Get a life.

Re: Karpathy’s Pelican

#437
post #411

Earlier quoted context omitted.

So if we test models that only output text to directly generate waveforms of sound from text or binary code, or directly generating binary code to replace a compiler, does that mean it is "intelligent"? Does that even test for intelligence? This is like testing if a horse can fly just because someone showed an image of a Pegasus, or testing if a fish can climb up a tree and believing they are not intelligent because…

I'm not sure where you are coming from, but we are testing a model that can create images to create an image for us. If it can't even do that well then I'm not sure why we need to talk about horses here. Pulling things into the ridiculous isn't condusive to a good faith discussion and does not help proving pseudoscientificness which you seem to be after. (I don't belive that the pelican bicycle test is meant to be a…

> I'm not sure where you are coming from, but we are testing a model that can create images to create an image for us.

Then use a model that supports multi-modal output then, instead of using models that only support raw text output to construct an image.

> If it can't even do that well then I'm not sure why we need to talk about horses here.

Testing a model that only outputs text and forcing it to output an unsupported modality is obviously not going to work well which is what it going on here.

Would you expect Claude/GPT to replace your compiler because it can output text? That would be like expecting that horses can fly because someone saw an ancient winged-horse on a brick tablet.

> Pulling things into the ridiculous isn't condusive to a good faith discussion and does not help proving pseudoscientificness which you seem to be after.

That is because the premise of the test is ridiculous and pseudoscientific.

> (I don't belive that the pelican bicycle test is meant to be a serious scientific endeavour, btw.)

Why did you say it was a "test for intelligence", when we both know it is a joke that tests for nothing?

Re: Karpathy’s Pelican

#439

Did I miss when they managed to draw a pelican? Cause all the ones I've seen are wrong in some way.

I think the SOTA is fable max. It's up to you whether that's good enough or not. Almost all other models do make some egregious mistakes with the bike. https://static.simonwillison.net/static/2026/fable-max.jpg

This is pretty close to "solved" I'd say. But then, it seems all the other models still need to catch up, and there's likely a little bit of room for improvement on Fable Max as well.

Re: Karpathy’s Pelican

#440
post #373

Did I miss when they managed to draw a pelican? Cause all the ones I've seen are wrong in some way.

Perfection is such an absolutely wild requirement for what we're seeing happen here, with the level of understanding required, from a tech that was complete fiction 5 years ago. Maybe excitement and wonder, in tech, is just something for us old guys, that watched it all be birthed. Get off my lawn!

The article makes it sound like Pelicans riding a bike are solved, and we can move on to something else. Nothing to do with excitement or wonder of the tech, it just claims one thing and I feel I've not seen that be the case so the claim seems incorrect to me.
Post reply on HN