Earlier quoted context omitted.
What do you mean by "oneshot"? The term applied to genai makes sense, but it doesn't make a whole lot of sense to apply the term to human art.
What doesn’t make sense? They aren’t drawing it by hand, they are still using SVG. Give them one chance before rendering it.
Karpathy’s Pelican
361–370 of 461 posts
Re: Karpathy’s Pelican
#362Re: Karpathy’s Pelican
#363Out of curiosity, I asked Opus 5 to do the same with the first ~1.5 pages of Neuromancer by William Gibson, which I thought might be a good test of interpretation from the model. It refused to use the text verbatim because of copyright (ironic), but the output was interesting nonetheless. https://claude.ai/public/artifacts/275dc3c2-7bd3-432b-94ff-d...
Also, large models refusing to work with copyright material is really hypocritical, copyright enforcement for thee but not for me
Re: Karpathy’s Pelican
#364Earlier quoted context omitted.
Trillions of dollars spent. Trillions of gigawatts consumed. And people still celebrate "Yay! We're less wrong than the other guys!" This is what the tech industry has become? Less of a failure is still failure.
I don't know how much money has been spent for AI, and I very much doubt you do. Do you know if more has been spent on LLM the last 9 years -- since "Attention is all you need" -- than Internet infrastructure during, say 1995 to 2004? That included the dotcom crash. Did you lament how a failure the Internet was? LLM has progressed a lot in the last two year, judging from the pelican drawings. I personally couldn't ca…
Pick up a newspaper. Start with the Wall Street Journal. These are public companies. It's not a secret.
Re: Karpathy’s Pelican
#365A lot of people are posting here about how bad the end product is, but that is kind of the point. Models have moved beyond generating images to a new kind of benchmark that better exposes understanding of the physical world, and we can use benchmarks like this to measure future progress. (Of course, it will have to be a qualitative/subjective measurement.)
But I do think this was a poor demonstration of the idea for another reason: Tolkien works have a HUGE corpus of training data. It's great that random users can come in and immediately recognize what the footage is, but it fails at the very first thing the pelican was meant to do:
- Render this thing you have only tangential training data of, that also happens to be an asymmetrical object so we can see how much you fuck up the details if you somehow flip the orientation half the time.
They should have used an obscure story, not "Most Studied Piece of Literally Work of The Past Century trademarksymbol"
Re: Karpathy’s Pelican
#366The difference between this and Simon's pelican is that with Simon, I get the prompt. Last I checked, I did not see the prompt for this really cool thing, so it is not reproducible. Did I miss the prompt somewhere?
he said the prompt was the first paragraph of LoTR, but he didn't mention a preamble this guy seems to have taken that idea and got something similar/better, so likely the prompt isn't too special https://x.com/Izkimar/status/2083819741643178208?s=20
This is a bigger difference to the pelican than simple reproducibility steps. The pelican is intentionally esoteric, and thus open ended. The LotR is mundane and has a "correct" answer, aka, copy the movie.
It makes it a really awful test of capabilities. The pelican isn't a slop test. This crap is.
Re: Karpathy’s Pelican
#367A lot of people are posting here about how bad the end product is, but that is kind of the point. Models have moved beyond generating images to a new kind of benchmark that better exposes understanding of the physical world, and we can use benchmarks like this to measure future progress. (Of course, it will have to be a qualitative/subjective measurement.)
I don't know I think it is charming in a way that is lacking in nearly everything else an LLM tries to do creatively. I have always preferred the result of getting an LLM to draw an svg or make a procedural animation like this to the uncanny hyper-realistic result of diffusion image/video generation.
Given that limitation, it is incredible what it can still accomplish. And when it falls short, it is often in a charming way, if you are open to seeing it that way. It reflects something like a child's understanding of the world, not entirely wrong, just incomplete.
The fluttering cubes that might have been bees or butterflies were my favorite part.