My biggest gripe is that he's comparing probabilistic models (LLMs) by a single sample. You wouldn't compare different random number generators by taking one sample from each and then concluding that generator 5 generates the highest numbers... Would be nicer to run the comparison with 10 images (or more) for each LLM and then average.
And by a sample that has become increasingly known as a benchmark. Newer training data will contain more articles like this one, which naturally improves the capabilities of an LLM to estimate what’s considered a good „pelican on a bike“.
The last six months in LLMs, illustrated by pelicans on bicycles
41–50 of 244 posts
Re: The last six months in LLMs, illustrated by pelicans on bicycles
#42Earlier quoted context omitted.
That seems to be a completely inappropriate use case? I would not hire a blind artist or a deaf musician.
Yeah, that's part of the point of this. Getting a state of the art text generating LLM to generate SVG illustrations is an inappropriate application of them. It's a fun way to deflate the hype. Sure, your new LLM may have cost XX million to train and beat all the others on the benchmarks, but when you ask it to draw a pelican on a bicycle it still outputs total junk.
https://chatgpt.com/share/684582a0-03cc-8006-b5b5-de51e5cd89...
Re: The last six months in LLMs, illustrated by pelicans on bicycles
#43Earlier quoted context omitted.
Ah yes the race to the bottom argument.
When I was at university, they got some people from industry to talk to us all about our CVs and how to do interviews. My CV had a stupid cliché, "committed to quality", which they correctly picked up on — "What do you mean?" one of them asked me, directly. I thought this meant I was focussed on being the best. He didn't like this answer. His example, blurred by 20 years of my imperfect human memory, was to ask me wh…
Go kart or porsche is irrelevant.
Re: The last six months in LLMs, illustrated by pelicans on bicycles
#44Earlier quoted context omitted.
I would not assume that this methodology applies to applied engineering, as a former actual real tangible meat space engineer. Things are a little nuanced and the nuances come from a combination of communication and experience, neither of which any LLM has any insight into at all. It's not out there on the internet to train it with and it's not even easy to put it into abstract terms which can be used as training dat…
I don't think any of that matters, CEOs will decide to use it anyway.
Re: The last six months in LLMs, illustrated by pelicans on bicycles
#45Re: The last six months in LLMs, illustrated by pelicans on bicycles
#46Here’s the spot where we see who’s TL;DR… > Claude 4 will rat you out to the feds! >If you expose it to evidence of malfeasance in your company, and you tell it it should act ethically, and you give it the ability to send email, it’ll rat you out.
Re: The last six months in LLMs, illustrated by pelicans on bicycles
#47My biggest gripe is that he's comparing probabilistic models (LLMs) by a single sample. You wouldn't compare different random number generators by taking one sample from each and then concluding that generator 5 generates the highest numbers... Would be nicer to run the comparison with 10 images (or more) for each LLM and then average.
It might not be 100% clear from the writing but this benchmark is mainly intended as a joke - I built a talk around it because it's a great way to make the last six months of model releases a lot more entertaining. I've been considering an expanded version of this where each model outputs ten images, then a vision model helps pick the "best" of those to represent that model in a further competition with other models.…
Any concerns about open source “AI celebrity talks” like yours can be used in contexts that would allow LLM models to optimize their market share in ways that we can’t imagine yet?
Your talk might influence the funding of AI startups.
#butterflyEffect
Re: The last six months in LLMs, illustrated by pelicans on bicycles
#48Earlier quoted context omitted.
I would not assume that this methodology applies to applied engineering, as a former actual real tangible meat space engineer. Things are a little nuanced and the nuances come from a combination of communication and experience, neither of which any LLM has any insight into at all. It's not out there on the internet to train it with and it's not even easy to put it into abstract terms which can be used as training dat…
https://www.solidworks.com/lp/evolve-your-design-workflows-a...
Re: The last six months in LLMs, illustrated by pelicans on bicycles
#49My biggest gripe is that he's comparing probabilistic models (LLMs) by a single sample. You wouldn't compare different random number generators by taking one sample from each and then concluding that generator 5 generates the highest numbers... Would be nicer to run the comparison with 10 images (or more) for each LLM and then average.
It might not be 100% clear from the writing but this benchmark is mainly intended as a joke - I built a talk around it because it's a great way to make the last six months of model releases a lot more entertaining. I've been considering an expanded version of this where each model outputs ten images, then a vision model helps pick the "best" of those to represent that model in a further competition with other models.…
Karpathy used it as an example in a recent interview: https://www.msn.com/en-in/health/other/ai-expert-asks-grok-3...
Re: The last six months in LLMs, illustrated by pelicans on bicycles
#50My biggest gripe is that he's comparing probabilistic models (LLMs) by a single sample. You wouldn't compare different random number generators by taking one sample from each and then concluding that generator 5 generates the highest numbers... Would be nicer to run the comparison with 10 images (or more) for each LLM and then average.
It might not be 100% clear from the writing but this benchmark is mainly intended as a joke - I built a talk around it because it's a great way to make the last six months of model releases a lot more entertaining. I've been considering an expanded version of this where each model outputs ten images, then a vision model helps pick the "best" of those to represent that model in a further competition with other models.…