I really enjoy Simon’s work in this space. I’ve read almost every blog post they’ve posted on this and I love seeing them poke and prod the models to see what pops out. The CLI tools are all very easy to use and complement each other nicely all without trying to do too much by themselves. And at the end of the day, it’s just so much fun to see someone else having so much fun. He’s like a kid in a candy store and that…
The last six months in LLMs, illustrated by pelicans on bicycles
91–100 of 244 posts
Re: The last six months in LLMs, illustrated by pelicans on bicycles
#92Earlier quoted context omitted.
It wasn't until I put these slides together that I realized quite how well my joke benchmark correlates with actual model performance - the "better" models genuinely do appear to draw better pelicans and I don't really understand why!
until they start targeting this benchmark
Re: The last six months in LLMs, illustrated by pelicans on bicycles
#93Earlier quoted context omitted.
You are right, but the companies making these models invest a lot of effort in marketing them as anything but probabilistic, i.e. making people think that these models work discretely like humans. In that case we'd expect a human with perfect drawing skills and perfect knowledge about bikes and birds to output such a simple drawing correctly 100% of the time. In any case, even if a model is probabilistic, if it had c…
> In that case we'd expect a human with perfect drawing skills and perfect knowledge about bikes and birds to output such a simple drawing correctly 100% of the time. Look upon these works, ye mighty, and despair: https://www.gianlucagimini.it/portfolio-item/velocipedia/
Re: The last six months in LLMs, illustrated by pelicans on bicycles
#94> This was one of the most successful product launches of all time. They signed up 100 million new user accounts in a week! They had a single hour where they signed up a million new accounts, as this thing kept on going viral again and again and again. Awkwardly, I never heard of it until now. I was aware that at some point they added ability to generate images to the app, but I never realized it was a major thing (p…
Except this went very mainstream. Lots of turn myself into a muppet, what is the human equivalent for my dog, etc. TikTok is all over this. It really is incredible.
Re: The last six months in LLMs, illustrated by pelicans on bicycles
#95Earlier quoted context omitted.
I'd say definitely do not do that. That would make the benchmark look more serious while still being problematic for knowledge cutoff reasons. Your prompt has become popular even outside your blog, so the odds of some SVG pelicans on bicycles making it into the training data have been going up and up. Karpathy used it as an example in a recent interview: https://www.msn.com/en-in/health/other/ai-expert-asks-grok-3...
I would definitely say he had no intention of doing that and was doubling down on the original joke.
clarification: I enjoyed the pelican on a bike and don't think it's that bad =p
Re: The last six months in LLMs, illustrated by pelicans on bicycles
#96Earlier quoted context omitted.
Except this went very mainstream. Lots of turn myself into a muppet, what is the human equivalent for my dog, etc. TikTok is all over this. It really is incredible.
The big trend was around the ghiblification of images. Those images were everywhere for a period of time.
Re: The last six months in LLMs, illustrated by pelicans on bicycles
#97Earlier quoted context omitted.
I'd say definitely do not do that. That would make the benchmark look more serious while still being problematic for knowledge cutoff reasons. Your prompt has become popular even outside your blog, so the odds of some SVG pelicans on bicycles making it into the training data have been going up and up. Karpathy used it as an example in a recent interview: https://www.msn.com/en-in/health/other/ai-expert-asks-grok-3...
I would definitely say he had no intention of doing that and was doubling down on the original joke.
Re: The last six months in LLMs, illustrated by pelicans on bicycles
#98Earlier quoted context omitted.
I'd say definitely do not do that. That would make the benchmark look more serious while still being problematic for knowledge cutoff reasons. Your prompt has become popular even outside your blog, so the odds of some SVG pelicans on bicycles making it into the training data have been going up and up. Karpathy used it as an example in a recent interview: https://www.msn.com/en-in/health/other/ai-expert-asks-grok-3...
Yeah, this is the problem with benchmarks where the questions/problems are public. They're valuable for some months, until it bleeds into the training set. I'm certain a lot of the "improvements" we're seeing are just benchmarks leaking into the training set.
The number of subject-verb-objects are near infinite. All are imaginable, but most are not plausible. A plausibility machine (LLM) will struggle with the implausible, until it can abstract well.
Re: The last six months in LLMs, illustrated by pelicans on bicycles
#99Re: The last six months in LLMs, illustrated by pelicans on bicycles
#100But bicycles are famously hard for artists as well. Cyclists can identify all of the parts, but if you don't ride a lot it can be surprisingly difficult to get all of the major bits of geometry right.