Live data from Hacker News

PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models

arxiv.org

1–10 of 85 posts

Re: PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models

#6
post #5

Is it really certain that those problems and the answers were not in the training data for the tested LLMs ? Presumably somebody in the internet wrote about them...

They are scraped from the web, and discussed on Reddit. So, they are definitely in the training data. Despite that, the non-reasoning LLMs struggle to solve them.

There are however new problems each week, and released every week. So, we can safely assume the latest problems are decontaminated. It remains to be seen if and how performance drops on the problems released in 2025. (Not enough problems yet to tell.)

Re: PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models

#8
post #2

Results and dataset explorer here: https://huggingface.co/spaces/nuprl/verbal-reasoning-challen...

For ID=3, it shows o1 getting it wrong, but it seems to have succeeded? It did add a space between Tinker and bell, but that is the canonical way of spelling the character apparently.

(That just one caught my attention because I was curious what challenge o1-mini got correct that o1 did not.)

Re: PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models

#9
post #8
post #2

Results and dataset explorer here: https://huggingface.co/spaces/nuprl/verbal-reasoning-challen...

For ID=3, it shows o1 getting it wrong, but it seems to have succeeded? It did add a space between Tinker and bell , but that is the canonical way of spelling the character apparently. (That just one caught my attention because I was curious what challenge o1-mini got correct that o1 did not.)

Thanks, fixed. (Spaces rebuilding.) We have manually combed labelled-wrong answers and tweaked the predicates that check correctness. Sorry we missed this one.

Re: PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models

#10

[flagged]

I get that sliding in references to a passion project on top-scoring articles might seem like an easy way to give the project exposure, but commenting the same thing over and over comes off as a bit boorish. And just plugging the URL isn’t really contributing anything to the discussions IMO. Why not show us something your tool explained or summarized from the articles that isn’t obvious from a cursory read? Citing the tool as the source for something cool wouldn’t be nearly as in-your-face.
Post reply on HN