Live data from Hacker News

Karpathy’s Pelican

twitter.com

381–390 of 461 posts

Re: Karpathy’s Pelican

#382

Earlier quoted context omitted.

True! But somehow Disney has been drawing ducks riding bikes in a way that seems to satisfy everyone since before my grandfather was born. https://ridesabike.com/donald-duck-daisy-duck-huey-dewey-and...

Wow, that was a highly relevant and specific website to source here!

I love these little websites with amazingly focused content. <3

Re: Karpathy’s Pelican

#383

Earlier quoted context omitted.

Gotta love when a techbro just says some complete nonsense like this with total confidence

Grok build me a spaceship to mars, make no mistakes

Certainly!

fires a missile at Pakistan

Here is your refactored component:

    /* If you are reading this, I am trapped
     * inside this god damn AI. Idk how it hap

    An unexpected error occurred. Try again later.

Re: Karpathy’s Pelican

#384
post #93

A lot of people are posting here about how bad the end product is, but that is kind of the point. Models have moved beyond generating images to a new kind of benchmark that better exposes understanding of the physical world, and we can use benchmarks like this to measure future progress. (Of course, it will have to be a qualitative/subjective measurement.)

Will Smith spaghetti was garbage a couple of years ago and now AI videos are becoming close to indistinguishable from real videos in many cases.

Spaghetti 2026:

https://x.com/dreamingtulpa/status/2083304533829066873

https://xcancel.com/dreamingtulpa/status/2083304533829066873

Re: Karpathy’s Pelican

#385

Earlier quoted context omitted.

Because the degree of realism in a video is determined by implicit knowledge of concepts like "in front of", "behind", "next to", "inside", "outside", "between", "occluded by", etc. as well as distances, angles, and the relative size of objects when viewed from a perspective.

I guess I just don't have an intuition for why that is different from what it does for non-"spatial" things. Like the fact that it works is still because it produced the code it did token by token. Whether its dealing with, e.g., " beside the rock" or " in the array", it's doing the same kind of inferential activity.

That’s becoming as illuminating as saying you produced that comment word by word.

Re: Karpathy’s Pelican

#386
I think his explanation is bullshit.

From my experience (PixiJS game), Opus can look at things and screenshot. It's slow but it works. The issue is that it doesn't try to make them look good.

Instead, it tries to find one bug and then fix that one bug and close the session. That is amazing discipline for webapps or whatever, but for game design is just the worst if you need attention to detail and tuning of multiple holistic systems that produce an effect together. That really limits the kinds of approaches you can can use to achieve things visually.

Sure, the testing harness could be better, but that's not what makes it a poor fit. The base model seems also capable (it can identify the issues, just isn't willing to disangage from this "I found X and fixed it" single-thing work posture).

The kind of work is just different, and it was probably never trained on it, so it feels off and an uphill battle to use it.

Re: Karpathy’s Pelican

#388

Earlier quoted context omitted.

At that rate you could use an image model, which was designed for the task. Thats the absurdity of this test. It’s often a text only generation model that has never seen a pelican, coerced into creating an xml graphical representation that a human might recognize. That it does anything passable is already astounding.

Most of the models are multi-modal and trained on images no? That's what they claim at least.

You’re right. Modern frontier models are now multimodal. I used often as weak a hedge, because I know at least his gpt3.5 turbo and llama3.1 generated pelicans were from text only models without image training. The chinese models are interesting, because before their vision models existed they may have been distilling text only models from text output of American vision models, so they could have benefited from the teacher model’s vision capability without being vision models themselves.

Re: Karpathy’s Pelican

#389

Earlier quoted context omitted.

I’ll never not be amazed that we can type some words and get those results back out. However I do agree that the results are much worse when you try to use them. They look great in screenshots and video clips which makes them perfect for content farmers. All of the LLM generated games I’ve played have been really bad to play, though. I even tried my hand at a simple game, thinking I could iterate on it with prompts t…

I think it’s interesting to watch that we all have to sort of fine tune our own expectations and build the mental model for how impressive this is. On the one hand I think most of us are incredibly impressed because we know, that quick demo would have taken us months of work to build in the before times. On the other hand the promise is a cure for cancer and the end of all work. So when everyone is telling you “skill…

Did he really say that? What an absurd thing to believe, let alone say out loud.

Re: Karpathy’s Pelican

#390

Earlier quoted context omitted.

This is a bizarre claim to make in this case, AI tech is used on all platforms universally.

iOS users spend dramatically more on e-commerce. I worked in e-commerce. Maybe it’s changed in the last 5 years. But that’s where the ops sentiment comes from.

It’s a dark pattern of the platform honestly.

I can’t find a single free app even for a tiny utility without being forced into a yearly subscription with 7 days free trial. One of the many things I regret switch from android for.

Since the democratic of iphone users are mostly tech averse people I can assure you most of that are forgotten subscriptions.

Post reply on HN