PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models
1–10 of 85 posts
Re: PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models
#2Re: PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models
#3Re: PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models
#4really great work! are you a co-author?
Re: PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models
#5Re: PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models
#6Is it really certain that those problems and the answers were not in the training data for the tested LLMs ? Presumably somebody in the internet wrote about them...
There are however new problems each week, and released every week. So, we can safely assume the latest problems are decontaminated. It remains to be seen if and how performance drops on the problems released in 2025. (Not enough problems yet to tell.)
Re: PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models
#7Re: PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models
#8Results and dataset explorer here: https://huggingface.co/spaces/nuprl/verbal-reasoning-challen...
(That just one caught my attention because I was curious what challenge o1-mini got correct that o1 did not.)
Re: PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models
#9Results and dataset explorer here: https://huggingface.co/spaces/nuprl/verbal-reasoning-challen...
For ID=3, it shows o1 getting it wrong, but it seems to have succeeded? It did add a space between Tinker and bell , but that is the canonical way of spelling the character apparently. (That just one caught my attention because I was curious what challenge o1-mini got correct that o1 did not.)
Re: PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models
#10[flagged]