Live data from Hacker News

Karpathy’s Pelican

twitter.com

401–410 of 461 posts

Re: Karpathy’s Pelican

#401

A simple prompt that still stumps frontier LLMs most of the time is “create a pinball game”. They’ll put all the right pieces there but then fail to arrange them such that the game is truly playable. They’ll put a wall in the way of the launch chute so the ball can’t be launched. Or the flippers will pivot the wrong way. Or there will be holes such that the ball drops off the bottom without getting within reach of th…

Failed demos like this give weight to the argument that AIs need more of a world model, an understanding of how physics works to avoid obvious stumbles like this.

Re: Karpathy’s Pelican

#402

Did I miss when they managed to draw a pelican? Cause all the ones I've seen are wrong in some way.

I think the SOTA is fable max. It's up to you whether that's good enough or not. Almost all other models do make some egregious mistakes with the bike.

https://static.simonwillison.net/static/2026/fable-max.jpg

Re: Karpathy’s Pelican

#403

> I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom This is an odd take, given that Karpathy is certainly aware that the LotR films absolutely did create Bag End in digital format; that their creation was outstandingly high quality; and that Claude’s output here very obviously “leans heavily” on their prior art.

Not sure whether Karpathy meant it like that or thought about this statement being interpreted this way, but I totally agree with you. Creating animated movies and special effects is placing a lot of 3d polygons by hand. Maybe not with JavaScript though. However, people also made things like

telnet towel.blinkenlights.nl

Now, were they „in their right mind“? I don‘t know. The more likely explanation is they found joy in it. Reminds me of what the Suno CEO Mikey Schulman said about making music: „[…] I think the majority of people don’t enjoy the majority of time they spend making music.“ (https://news.ycombinator.com/item?id=42688538)

Re: Karpathy’s Pelican

#405

Earlier quoted context omitted.

iOS users spend dramatically more on e-commerce. I worked in e-commerce. Maybe it’s changed in the last 5 years. But that’s where the ops sentiment comes from.

It’s a dark pattern of the platform honestly. I can’t find a single free app even for a tiny utility without being forced into a yearly subscription with 7 days free trial. One of the many things I regret switch from android for. Since the democratic of iphone users are mostly tech averse people I can assure you most of that are forgotten subscriptions.

Having to pay for stuff you use is not a dark pattern.

Re: Karpathy’s Pelican

#406
post #241

Earlier quoted context omitted.

The mean human gets a middling score in creative writing tasks at the end of their mandatory education, and most then leave school and forget what little they ever learned outside whatever their career path happened to be. Most humans never do a creative writing course after school, and the longest fiction most people will write is their resume description of what their previous jobs involved, or perhaps their dating…

This is so absurdly condescending and simultaneously incorrect that it astounds me you’re able to function as part of a society.

> it astounds me

This should cause you to reconsider your beliefs leading up to it.

GPT-4 (!) has, in studies comparing multiple models with humans for creativity, beaten the mean human. In comparison, even just the mean of the top 50% of humans beat the models studied in that case (link follows), but the point is that if you think the mean humans is particularly noteworthy, you've avoided the half of the population who have the creativity of a pot noodle.

https://www.nature.com/articles/s41598-025-25157-3/figures/2

There's more studies out there with other types of creativity test, but the conclusion is basically the same: the best models beat the mean human, but are nowhere near as good the worst *publishable* human.

My guess is this is both why LLM slop happens and why it grates so hard: a significant number of bosses look at the output and think to themselves "wow so creative" because it's more creative than they themselves are; but those bosses weren't hired to be creative, they were hired to be a boss, and all the people who they hired to be creative are going "arg, no, can't you see how bad this is?"... but that is just a guess, I've not found any surveys comparing *management* creativity to LLMs, closest is e.g. this about decision making, not creativity: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5156585

Re: Karpathy’s Pelican

#407

It seems pretty clear that Anthropic models have been specifically trained to be good at generating three.js (JavaScript 3-D Graphics) code, so given current state of AI code generation in general, I don't find three.js models/animations as indicative of anything other than the model's ability to write three.js code. When Fable was first released the day-1 demos of it on Twitter (presumably from people who were given…

Taking a single paragraph of literary text, which is abstract and ambiguous, and converting it into a 3D animation requires an enormous amount of implicit knowledge about spatial relationships, intuitive physics, everyday objects, and so forth. Not to mention the mathematics of 3D transformations and computer graphics more generally. Saying that it's indicative of no more than three.js coding ability is absurd.

*Of text it has read the full book of, read a huge corpus of text discussing it and watched the highly successful adaption of.

We should feed a snippet of an unreleased book in a novel universe.

Re: Karpathy’s Pelican

#408

It seems pretty clear that Anthropic models have been specifically trained to be good at generating three.js (JavaScript 3-D Graphics) code, so given current state of AI code generation in general, I don't find three.js models/animations as indicative of anything other than the model's ability to write three.js code. When Fable was first released the day-1 demos of it on Twitter (presumably from people who were given…

[flagged]

The guidelines ask:

Please don't fulminate. Please don't sneer, including at the rest of the community. https://news.ycombinator.com/newsguidelines.html

Many people from all sections of HN, the tech industry and broader society have been surprised and wrong in all directions about how the emergence of AI is playing out. I certainly have. I can't think of a single person whose predictions have been precisely accurate. So, please don't use terms like “wall of shame” and “cope”. It doesn't help anyone make better predictions and only makes you, and HN generally, seem mean.

Re: Karpathy’s Pelican

#409
We've been building motion graphics capability (on chatoctopus.com), and "closing the loop" has been one of the hardest challenges. Coding agents have it easy; static analysis, lints, unit tests, etc.

But when visual perception and "taste" get involved, it becomes a lot harder.

Re: Karpathy’s Pelican

#410

Earlier quoted context omitted.

I yearn for the old days of bad photoshop/wordart/powerpoint flyers. They were visually bad but honest and sometimes soulful or playful. The AI slop that's everywhere now looks superficially more professional but it's very busy, samey, unnatural and it really turns me off.

I don't want to say this is indicative, but my local library used AI-generated event poster/images for the past year. But recently they switched back to word/clipart event posters. When I was checking out, I query the librarian about it, and apparently they saw a 40-60% drop in attendance. On one hand, I found that hard to believe, but on the other hand, simple word/clipart event posters do have a certain inviting ch…

I think that word/clipart posters make it clearer the date and time, with GenAI material you end upo adding a big colorful image that takes most of the space and relegate the important bit to a second place. At least that's what I see around me.
Post reply on HN