I think one main failure in the framing of these papers (and discussion of LLMs more broadly) is that the abstract says that GPT4 ‘struggles’ with logical reasoning: > ChatGPT and GPT-4 do relatively well on well-known datasets […] however, the performance drops significantly when handling newly released and out-of-distribution [where] Logical reasoning remains challenging for ChatGPT and GPT-4 But reading the paper…
If you really force it to reason, rather than regurgitate arguments from its training set, you will find it is nowhere near the genius line. Make up some rules and have it try to answer questions according to the rules. In my experiments I feel it's something like a 4 or 5 year old child both in its logical limitations and penchant for distraction. However it's important to note one VERY important thing -- this is no…
For what it's worth, neither are we, really. Not disagreeing with anything you're saying, just musing.
> superhuman levels of reasoning
This one has always stumped me a bit though. I'm not quite sure what that looks like. Laplace's Demon?