Earlier quoted context omitted.
> I think it's touching the limit of what one can reasonably expect any intelligent thing to produce with the only direction being "produce an svg of a pelican riding a bicycle". Oh come now. I am extremely confident that if I hired a professional artist to draw a picture of a pelican riding a bicycle, I would get something inarguably much better than what today's best coding LLMs can produce.
> if I hired a professional artist to draw a picture of a pelican riding a bicycle, I think AI folks have done a terrible job of communicating this, but replacing a professional simply isn't the point. The point is to serve all the situations where people would've never considered hiring a professional, and where perfection or artistic merit isn't the point (say a personal throwaway recreation of an LOTR world). And…
Karpathy’s Pelican
351–360 of 461 posts
Re: Karpathy’s Pelican
#352Re: Karpathy’s Pelican
#353The difference between this and Simon's pelican is that with Simon, I get the prompt. Last I checked, I did not see the prompt for this really cool thing, so it is not reproducible. Did I miss the prompt somewhere?
he said the prompt was the first paragraph of LoTR, but he didn't mention a preamble this guy seems to have taken that idea and got something similar/better, so likely the prompt isn't too special https://x.com/Izkimar/status/2083819741643178208?s=20
In either case, I still don't see how I could reproduce this to test against various models, which is the entire point of Simon's pelican.
Re: Karpathy’s Pelican
#354Earlier quoted context omitted.
Agree. The pelican benchmark was interesting a year ago when most models struggled and a good pelican indicated an unusually capable model. Now it’s saturated and uninteresting. A good new benchmark should have awful performance to start and there should be a lot of headroom for improvement. This benchmark is also intentionally difficult and requires the LLM to develop the animation through spatial reasoning and firs…
Have you seen pelicans in Simon Willison’s tests? It is still not a pelican on bike I would like to publish :) Imho we don’t need to make benchmarks that draw the whole 3D world. Pelican’s drawing is really nice in its simplicity and complexity at the same time. It seems to be an obscene waste of compute time to generate useless 3D worlds that are just a bragging - 3D is really heavy discipline to make it right, see…
No, this demo is the useless 3D world, and you're bragging.
A real game would have a lot more immersive of a world, and you wouldn't need to.
Re: Karpathy’s Pelican
#355It refused to use the text verbatim because of copyright (ironic), but the output was interesting nonetheless.
https://claude.ai/public/artifacts/275dc3c2-7bd3-432b-94ff-d...
Re: Karpathy’s Pelican
#356I think it's useful when evaluating "model+intended harness", but I'm more interested in seeing raw model improvements than harness improvements.
Re: Karpathy’s Pelican
#357Re: Karpathy’s Pelican
#358It seems pretty clear that Anthropic models have been specifically trained to be good at generating three.js (JavaScript 3-D Graphics) code, so given current state of AI code generation in general, I don't find three.js models/animations as indicative of anything other than the model's ability to write three.js code. When Fable was first released the day-1 demos of it on Twitter (presumably from people who were given…
Why do you think it's pretty clear? Anthropic models are excellent at working with Blender APIs, other industry standard 3D modelling programs and tasks, and game libraries that have nothing to do with Three.js. Three.js is a quite popular library; and browser-based apps are more easily sharable and more portable. So models having a preference for using it when unprompted doesn't suggest anything, just like how model…
It seems they trained it to be good at it, then requested everyone to demo it.
Re: Karpathy’s Pelican
#359Re: Karpathy’s Pelican
#360IMO, the area where AI is going to be most useful over the next couple years is in developing manufacturing processes top to bottom. Maybe a million token budget is too small, but something like "design me a sneaker and all the equipment to manufacture it autonomously".
I've seen them make time estimates that are physically impossible, expect medications to be in two places at once, etc.