There was an interesting thought from a Microsoft researcher in an episode of This American Life. He had been given early access to GPT4 and found that it had gained an understanding of gravity and balancing. 3.5 would fail to describe a safe order for stacking 3 eggs, a bottle, a book and a nail. But 4 would give robust answers with logical justification added to each construction step for the tower. The researcher…
> Give the same task to GPT4 and it proves it’s modeled the concept appropriately because it can pull the right words out of the ether. I take your point, but we should be careful when using phrased like “modeled the concept appropriately ” here. That implies correctness. There are many ‘concept models’ that could work well enough. Unless we can inspect the concept model directly (inside the LLM, somehow!) we are lef…
All models are wrong, but some are useful.
Perhaps we should say "modeled the concept usefully".
> In some ways, my standard here is probably not that different than a rigorous human educational assessment. I’m not talking about standardized tests, I’m talking about adversarial challenges like one would hope to see during a Ph.D. defense.
Usefulness is very contextual obviously. Newton's laws aren't entirely correct but are still useful.