There was an interesting thought from a Microsoft researcher in an episode of This American Life. He had been given early access to GPT4 and found that it had gained an understanding of gravity and balancing. 3.5 would fail to describe a safe order for stacking 3 eggs, a bottle, a book and a nail. But 4 would give robust answers with logical justification added to each construction step for the tower. The researcher…
I take your point, but we should be careful when using phrased like “modeled the concept appropriately” here. That implies correctness.
There are many ‘concept models’ that could work well enough. Unless we can inspect the concept model directly (inside the LLM, somehow!) we are left to reason about how the concept is represented via statistics over a set of interactions. So, if, over the course of many interactions, it seems that a model is saying reasonable things we can have more confidence that it has achieved a concept model that is “good enough”. But such an assessment is always in relationship to the test protocol. Claims to generality should be done carefully.
In some ways, my standard here is probably not that different than a rigorous human educational assessment. I’m not talking about standardized tests, I’m talking about adversarial challenges like one would hope to see during a Ph.D. defense.