This comment and the one above from astrobiased feel like coming into a messy codebase, and it’s more work to sort it out than it would have been to write it from scratch.... And since I actually do intelligence testing as a clinical psychologist, I have experience with this in both practice and theory. So now I’m going to waste an hour because I just have to respond to “something is wrong on the internet.”.
Chollet's distinction is useful. High performance on known tasks is not the same thing as efficient adaptation to a novel task. Prior knowledge and training data can buy skill. That is a central point of On the Measure of Intelligence. But it does not follow that current frontier progress is only "coverage-driven competence." That is a hypothesis. It is not a result established by Chollet's framework.
"Overfitting at scale" is also the wrong term. A model that learns broad representations and applies them successfully to unseen examples is generalizing. The relevant concern is whether apparent novelty is actually inside the effective training distribution, not whether the model is "overfit."
There is also an unstated premise here: that adding broad knowledge and skills cannot improve the machinery used for novel problem solving. I do not see a basis for assuming that. Learned representations, abstractions, reasoning patterns, and cross-domain analogies can themselves support transfer to new tasks. Whether this becomes sufficient for general intelligence is an open question with insufficient data. But its a perfectly valid hypothesis right now that, given enough domain knowledge and symbolic reasoning examples, LLM COULD maybe "Grok" AGI at a certain critical threshold.
And ARC-AGI-3 was specifically designed around novel abstract environments that require exploration and adaptation. Astra scores 99.9% with OpenAI's context-preserving Provider Adapter, and ARC reports that Astra constructed compact symbolic models of unfamiliar environments. That does not prove AGI, but it points in that direction more so than the other way around.
Gc roughly maps to acquired knowledge. Gf roughly maps to reasoning in relatively novel situations. Naming those two categories does not tell us whether increasing acquired knowledge and learned abstractions in an AI can improve Gf-like behavior. That causal question is exactly what is disputed.
And "Frontier models probably have maxed out crystallized intelligence" is just obviously wrong, unless you think they have been able to dig up every a scrap of paper with knowledge/information on it in the entire world, AND that there is no more useful knowledge to be generated left in the universe.
And the statement that intelligence and creativity are independent is simply wrong. A meta-analysis of 112 studies and 34k participants found a positive correlation of about r .25 between intelligence and divergent thinking. It also found that using g, Gf, or Gc did not eliminate that relationship. Creative achievement has a smaller but still positive meta-analytic association with intelligence, around r = .16. These are distinct constructs, not independent constructs.
And this is just a bad take: "AIs are terrible at creativity". At best that depends on which creativity, and I think its straight up wrong. On divergent thinking tasks, the operationalization behind every ADHD study you could cite, LLMs score above most humans, with the top humans still ahead. If you means Big-C, paradigm-shifting creativity, that is a different construct and none of the ADHD evidence transfers to it.
And if I where to say what I subjectively feel and see.... I have ABSOLUTELY no idea how people can say that we are not seeing sparks of creativity from AIs already. If a PERSON produced some of the music, solutions or deductions that I have seen AIs do, people would have NO problem celebrating it as extremely creative.
And finally, the ADHD claim is also, at best, overstated and just as often debunked. There is some evidence that higher subclinical ADHD trait scores, often survey studies only, are associated with better performance on some divergent-thinking measures. But a review of 31 studies did not find a consistent creativity advantage for people with clinical ADHD, and it found no evidence of better convergent thinking.
Okay, I’m done… And nobody noticed that I’m not doing my job here.