Live data from Hacker News

Karpathy’s Pelican

twitter.com

241–250 of 461 posts

Re: Karpathy’s Pelican

#241
post #132

Earlier quoted context omitted.

While they are bad at creative work*, this reasoning isn't going to show it. If you took the best, most creative, human writer in the world, and for thought experiment reasons they had amnesia (to mimic AI blank context windows) specifically while you asked them for a story idea 100 times in a row, my expectation is that this human would also give you the same idea at least 80 times out of that 100. * still better th…

Better than the mean human? The mean human creates far better stories while daydreaming. AI enthusiasts have such a distorted view of human capability, it’s bizarre.

The mean human gets a middling score in creative writing tasks at the end of their mandatory education, and most then leave school and forget what little they ever learned outside whatever their career path happened to be.

Most humans never do a creative writing course after school, and the longest fiction most people will write is their resume description of what their previous jobs involved, or perhaps their dating profile.

Don't mistake what you see published (or what your friends are like) for the average human: the average Hacker News comment easily above the writing grade (and creativity) of e.g. many of the one-shots stories I've seen attempted on some creative writing subreddits. "Mean" is not a high bar.

Re: Karpathy’s Pelican

#242

Earlier quoted context omitted.

This is a bizarre claim to make in this case, AI tech is used on all platforms universally.

Literally just google “mac users more likely to pay” and you will find many such instances

Are Mac users more likely to pay? Likely yes.

Are the total number of Mac users who pay higher than the total number of Windows users who pay? I kind of doubt it. Windows market share is still much higher than MacOS.

Re: Karpathy’s Pelican

#243
post #99

I really dislike this AI programming thing of “Mr LLM, go slam your face into the problem until there’s no problem left, then call me back”. (Not sure if it’s a recent trend or a fundamental nature.) It always brings to my mind some words from Rich Hickey: I think we’re in this world I’d like to call “guardrail programming”. It’s really sad: we’re like, “I can make change because I have tests!”. Who does that? Who dr…

To me, that's a totally normal way of using computers. Repeating tasks is what computers do best, AI or not.

Take chess engines of the "deep blue" era for instance. These engines are stupid, trying millions of moves that are obviously terrible, no human chess player would do that. And yet, this is what worked best, because computers are so good at repetition. Recent development using neural networks made chess engines smarter, but using repetition is still how the beat humans.

What you call "guardrail programming" in another context would genetic algorithms, a technique that has recognized applications. And if you look at videos of genetic algorithms learning to play racing games, it literally looks like a bunch of carnival bumper cars onto the highways, but done well, after some generations, it becomes competent driving, sometimes even record breaking.

It can be expensive, it is often the case for LLMs, so you may want to try to be a bit smarter at first to spare some resources, but to me, it is just using computers as intended.

Re: Karpathy’s Pelican

#244
post #202

I’d like to see a human one shot a pelican on a bicycle in raw svg.

It's been shown in tests that most humans can not from memory draw a functional bicycle given pen and paper.

Everyone knows roughly what a bicycle looks like - wheels, frame, seat, peddles, handlebars etc, but the details of the frame and exactly how the other parts connect to it throws people off. It seems people memorize the "concept" of a frame, but not the specifics.

Try it without cheating, then google a picture of a bicycle!

Even if you have memorized what a bicycle (and pelican) look like, writing code to draw one is not the sort of thing humans are good at any more than they are good at mentally calculating cube roots - better to use a computer for stuff like that.

Re: Karpathy’s Pelican

#245
post #8

I worked with an LLM to build a ~3D animation of the Back to the Future delorean Time Machine as a way to spice up the hero on a docs page. That took a fair amount of custom tuning and I had to create a tuning view to get some of the behaviors right. But it was enough fun that I generalized it to take in ~any scene description from a film. It goes out and gets more detailed descriptions and film stills if available b…

Ha, that's fun! I've built half your product for myself as a personal, native MacOS app, with a slightly different direction, but the same core idea. I might go do a little scavenging for ideas in your docs :-)

Re: Karpathy’s Pelican

#246

IMO, the area where AI is going to be most useful over the next couple years is in developing manufacturing processes top to bottom. Maybe a million token budget is too small, but something like "design me a sneaker and all the equipment to manufacture it autonomously".

They can't even run a vending machine. I don't want to do a shallow dismissal, but I think there's a gulf between my understanding and yours. I hope it's me so I learn something. I think LLMs will be excellent glue of "find the right function/button and run/push it" but design without constraints and they just explode immediately

I've been getting into 3D printing recently and while the flagship Anthropic models can regurgitate community wisdom about the hobby they can't do spatial reasoning. They'll tell me that objects will fit on my print bed when they won't and vice-versa. They get confused about materials and nozzles. It takes several iterations to get them to do basic objects (in my case a slotted tray) in OpenSCAD. Useful, certainly, but very far from autonomous for end-to-end manufacturing workflows.

Re: Karpathy’s Pelican

#247
post #165
post #158

Earlier quoted context omitted.

Aren't they still bad at understanding how bicycle frame works? Especially the steering part?

Maybe that makes them human.. https://www.booooooom.com/2016/05/09/bicycles-built-based-on...

I've fed a couple of those chicken-scratch sketches to nanobanana with prompt "Treat attached as a technical drawing of a bicycle. Produce photo-realistic image of the actual bike manufactured to that spec. Try to stay close to the input, where possible", and... wow, results are not great.

Re: Karpathy’s Pelican

#248

I don't think it's a bad way to benchmark new models, I just find it concerning that the author implies that "pelican on a bicycle" has been exhausted. At the risk of making overly broad, unfalsifiable claims I think multi-year exposure to AI content has dramatically raised our expectations for speed and volume but lowered them for quality. We see a very janky pelican and declare the problem solved.

Humans are drawing pelicans riding bicycles now. Just google it and you will find 5 or 10 of them in the first few results. Including a t-shirt design.

So it's a pretty much pointless test now.

Re: Karpathy’s Pelican

#249
I love how nobody cares about copyright anymore. Not even an after thought. Might be okay if you're an employee of anthropic/openai but for us mere mortals I'm not sure I would share something this blatant. In US it's $150k+ per violation and you've given them all the proof (even a confession)

Re: Karpathy’s Pelican

#250
> sure, why not, it's ~free

Yeah, please check in with the folks protesting data center builds in their town causing their electricity prices to skyrocket and tap water to turn into a scarce resource.

And now, instead of actually doing the above, please go ahead and downvote me, because how dare he question those LLM games.

Post reply on HN