Earlier quoted context omitted.
Can someone explain how we arrived at the pelican test? Was there some actual theory behind why it's difficult to produce? Or did someone just think it up, discover it was consistently difficult, and now we just all know it's a good test?
I set it up as a joke, to make fun of all of the other benchmarks. To my surprise it ended up being a surprisingly good measure of the quality of the model for other tasks (up to a certain point at least), though I've never seen a convincing argument as to why. I gave a talk about it last year: https://simonwillison.net/2025/Jun/6/six-months-in-llms/ It should not be treated as a serious benchmark.
if it is indeed a good measure of the quality of the model (hint: it's not) then, logically, it should be taken seriously.
this is, sadly, a great example of the kind of doublethink the "AI" hypesters (yes - whether you like it or not simon - that is what you are now) are all too capable of.