Live data from Hacker News

Karpathy’s Pelican

twitter.com

301–310 of 461 posts

Re: Karpathy’s Pelican

#301
post #166
post #158

Earlier quoted context omitted.

Aren't they still bad at understanding how bicycle frame works? Especially the steering part?

Shhh...you'll alert the models :) Totally agree though, anyone with a vague understanding of how bikes works ignores the pelican because they know the bike is unrideable in the first place.

> Shhh...you'll alert the models :)

They are listening.

Re: Karpathy’s Pelican

#302

this is not a good benchmark for models, but it's great if you're optimizing for attention on twitter because video content and 3d animations perform best on social media. a real benchmark is instead running evals on your own traces, and building a cost/quality/speed profile for models based on real workloads. but it doesn't get you a shiny video you can post on twitter.

You can do both.

Re: Karpathy’s Pelican

#303
post #251

I don't think it's a bad way to benchmark new models, I just find it concerning that the author implies that "pelican on a bicycle" has been exhausted. At the risk of making overly broad, unfalsifiable claims I think multi-year exposure to AI content has dramatically raised our expectations for speed and volume but lowered them for quality. We see a very janky pelican and declare the problem solved.

Agreed. It's far from solved. Modern LLMs still generate pelican bike SVGs with obvious errors: * some omitted the bottom of the diamond which connects from the pedals to the rear wheel * some added an extra connection from the pedals to the front wheel, making it impossible to steer * none could align the head tube with the fork * none added a correct offset to the fork * none could generate the chain properly in a…

what's fable at, does anyone know?

Re: Karpathy’s Pelican

#304

Earlier quoted context omitted.

AI has deeply changed the way I think, feel and act around a computer. In the same way that dialing into the internet changed things for me. Since using ChatGPT the first time until now I have never cared once to look at these pelicans on bikes people seem to get hung up about. It could never have been a thing and nothing would change. See the forest through the trees.

What you’re saying is that you’re not interested in benchmarks. But then you go a step further and state that this particular benchmark is entirely inconsequential. That’s like telling you that if you didn’t exist, nothing would change. Even if that were true, it would still be an insensitive and rude thing to say, wouldn’t it?

It's okay to be rude to benchmarks though, they don't have feelings.

Re: Karpathy’s Pelican

#306

Earlier quoted context omitted.

I think it's touching the limit of what one can reasonably expect any intelligent thing to produce with the only direction being "produce an svg of a pelican riding a bicycle". When you aren't sure if an LLM can write an svg well, or that it will be able to form a pelican shape, or animate a bicycle, it's a good test. After that, it's all judgement: how detailed should the pelican be? pelicans are the wrong shape for…

> I think it's touching the limit of what one can reasonably expect any intelligent thing to produce with the only direction being "produce an svg of a pelican riding a bicycle". Oh come now. I am extremely confident that if I hired a professional artist to draw a picture of a pelican riding a bicycle, I would get something inarguably much better than what today's best coding LLMs can produce.

> if I hired a professional artist to draw a picture of a pelican riding a bicycle,

I think AI folks have done a terrible job of communicating this, but replacing a professional simply isn't the point. The point is to serve all the situations where people would've never considered hiring a professional, and where perfection or artistic merit isn't the point (say a personal throwaway recreation of an LOTR world).

And I think in that regard the benchmarks are pretty good.

Re: Karpathy’s Pelican

#307
post #61

I'm pretty tired of the "Y made this game in Z tokens" all over the internet last week. They look impressive, and it's cool that it's even possible, but they're useless as games. None of them are any fun. They're like the most boring variant of basic controllers you can imagine. None have any cool mechanics. None have any tweaks made from hours and hours of testing. All have the same cel-shader.

LLMs are bad at creative work and I don’t see them improving any time soon. Try asking an agent to write a story about raccoons. It will almost certainly involve either stealing food or raiding trash with a 50% chance of having a character named Pip. If a location is mentioned, it’ll be Elm Street. I guess this is the average story and, similarly, the average game is boring and predictable

1/3

I did get pip, but it's a story about trying to rescue the moon from a lake and the raccoon lives on moonberry lane

What's pip from?

Re: Karpathy’s Pelican

#308

Earlier quoted context omitted.

> If the computer can't do it better than a human being, then what's the point? Because the benchmark wasn't testing "can an LLM draw a pelican like a human". The original article was testing the relative capabilities between LLMs. Now that LLMs can all draw pelicans all similarly, the test is less interesting as a comparative benchmark.

Trillions of dollars spent. Trillions of gigawatts consumed. And people still celebrate "Yay! We're less wrong than the other guys!" This is what the tech industry has become? Less of a failure is still failure.

That is the software industry. Our product is less bugged than our competitor’s.

Re: Karpathy’s Pelican

#309

Earlier quoted context omitted.

I don't think he's claiming it's been exhausted. It's just that things have progressed to a point where people are arguing over the finer points of which pelican looks better -- which is often a matter of taste, and an indication that we've hit the knee in benchmark where models are no longer failing in obviously awful ways.

I haven't really seen evidence that any ai can reliably draw a pelican riding a bicycle. Not if you look at the image long enough to take it in. Even the best ones have something wrong with them. Not a matter of taste but a matter of having both legs peddling on the viewer's side of the bicycle or having two beaks. I'm actually beginning to wonder if some people who ignore these things have a different, somewhat less…

Pelicans don't ride bicycles.

It's physically impossible.

The problem is to draw it in the least disturbing way possible.

Re: Karpathy’s Pelican

#310
post #270

Earlier quoted context omitted.

> I haven't really seen evidence that any ai can reliably draw a pelican riding a bicycle Please remember, we've started from there : https://simonwillison.net/2024/Oct/25/pelicans-on-a-bicycle/ When it started, it was clear what LLM would stand out, its style, etc. Nowadays, the pelicans look similar, the difference is in details and sometimes hard to catch. Sure, the task is not completed perfectly, but that's not…

Sure, the task is not completed perfectly, but that's not the point. Isn't it? If the computer can't do it better than a human being, then what's the point? Being wrong at scale is not better than being right.

Hugely profitable companies leak half the nation's personal data every month. Tell me more about how being wrong at scale is not valuable.
Post reply on HN