Live data from Hacker News

PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models

arxiv.org

11–20 of 85 posts

Re: PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models

#11
post #9
post #8

Earlier quoted context omitted.

For ID=3, it shows o1 getting it wrong, but it seems to have succeeded? It did add a space between Tinker and bell , but that is the canonical way of spelling the character apparently. (That just one caught my attention because I was curious what challenge o1-mini got correct that o1 did not.)

Thanks, fixed. (Spaces rebuilding.) We have manually combed labelled-wrong answers and tweaked the predicates that check correctness. Sorry we missed this one.

I thought maybe I was just missing something silly.. was hesitant to comment without additional coffee. And no problem, thanks for sharing this!

Re: PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models

#15
post #13

My feeling is that a lot of challenge could come from the tokenizer used by the model, similar to r in strawberry problem.

There is likely a little bit of that, but I don't think that's the primary factor that makes these hard. If you look at the reasoning performed by R1 ("Show Thoughts" checkbox), it has no problem splitting and joining candidate words into letters, or even doing the letter-level manipulations. Just to give an example:

Take a familiar brand name, seen along roads and highways in the United States. It has five letters, two syllables. The first syllable, phonetically, is a word that is the opposite of the word spelled by the second syllable. What brand name is it?

Okay, let me try to figure out this riddle. The question is about a brand name that's familiar on roads and highways in the US. It has five letters and two syllables. The key part is that the first syllable phonetically is the opposite of the second syllable spelled out. Hmm.

First, I need to think of common brands seen on highways. Gas stations, fast food, hotels, maybe car brands. Five letters, two syllables. Let's list some possibilities: Exxon, Shell, Chevron, Mobil, BP (but BP is two letters, so probably not). Maybe food places like Subway, but that's six letters. McDonald's is too long. Burger King? No. Maybe hotels: Motel 6? No, that's a number. Holiday Inn? Too long. Maybe Best Western? No, too long.

Wait, five letters. Let's think of two-syllable brands. Exxon (E-xxon), but Exxon is five letters. First syllable is "Ex", second is "xon". But what's the opposite of "xon"? That doesn't make sense. Maybe the first syllable is a word, and the second syllable when spelled is the opposite. So, for example, if the first syllable is "up", the second syllable's letters spell "down". But "updown" isn't a brand. Wait, the first syllable phonetically is a word, and the second syllable's spelling is the opposite. So when you say the first syllable, it's a word, and the letters of the second syllable spell its opposite.

So, for example, if the brand is "Updn" (made up), where "up" is the first syllable, and "dn" spells "down". But that's not a real brand. Let's think of real brands. Maybe "Comcast"? No, too long. "Verizon"? Also longer. "Sprint"? Five letters. Sprint. S-print. But "S" doesn't mean anything opposite of "print". Hmm. Maybe "Exxon" again. "Ex" as in former, and the opposite would be "xon" spelled as... no.

Wait, maybe the first syllable is a word that's the opposite of the word formed by the second syllable's letters. For instance, "Sunoco" is a gas station, but that's six letters. "Shell" is five letters but one syllable. "Hertz" is car rental. Five letters, one syllable. "Avis"? Four letters. "FedEx"? Five letters? F-E-D-E-X? No, that's six. Wait, FedEx is two syllables: Fed-Ex. Five letters? F-E-D-E-X? That's five letters? F, E, D, E, X. Yes, five letters. FedEx. The first syllable is "Fed", the second is "Ex". Now, "Fed" as in past tense of feed, but maybe "Fed" as in federal. The opposite of "Ex" could be "current" or "present". But that doesn't seem right. Alternatively, "Ex" as in former, so the opposite would be "current" or "new". Hmm, not sure. Maybe "FedEx" isn't the answer.

Another thought: "Amtrak" is a brand on highways, but that's six letters. "Greyhound"? Too long. "Dollar" car rental? Five letters? No, six. "Thrifty" is seven. Wait, maybe "Lyft"? No, four letters. "U-Haul" is hyphenated.

Wait, let's think of opposites. The first syllable's sound is a word, and the second syllable's spelling is its opposite. So, for example, if the brand is "Inote", where "I" (pronounced "eye") and "note", but "eye" and "note" don't have opposites. Alternatively, "Hi-Lo", but that's hyphenated. Maybe "Dayton" (a tire company?), but "Day" vs "ton" – not opposites.

Alternatively, maybe "Duracell" – no, too long. "Energizer"? No. "Eveready"? No. Wait, maybe "Goodyear". That's a tire company. Good-year. "Good" and "year". Not opposites.

Wait, the key is that the first syllable phonetically is a word, and the second syllable's spelling is the opposite. So, for instance, if the first syllable is "pro", the second syllable's letters spell "con". But "procon" isn't a brand. Or "up" and "down". Maybe "Updown", but not a real brand.

Wait, the brand "Exxon" – "Ex" and "xon". If "Ex" is the opposite of "xon" spelled as "xon" – no. Maybe "Ex" (as in former) and "xon" – but "xon" isn't a word. Alternatively, "Exxon" could be broken into "Ex" and "on". Wait, "Ex" (as in former) and "on" (the opposite being "off"). But "Exxon" is E-X-X-O-N, so the second syllable is "xon", not "on". Hmm.

Another approach: think of opposites like "yes/no", "hot/cold", "fast/slow", "high/low", "in/out", "on/off", "start/stop", "day/night". Maybe the brand incorporates one of these.

"Hi-Lo" is hyphenated. "Stop & Go" – no. "Speedpass" – not five letters. "7-Eleven" – no.

Wait, "Loews" hotels? Five letters. "Lo-ews". "Lo" as in low, and "ews" – not an opposite. Maybe not.

Wait, "Lucent" – a telecom company, but that's six letters. "Verizon" – seven letters.

Wait, "Motel 6" – no, includes a number. "Super 8" – same.

Wait, "Shell" – five letters but one syllable. "Chevron" – two syllables but seven letters.

Wait, "Pepsi" – five letters, two syllables. "Pep" and "si". "Pep" as in energy, and "si" – not an opposite.

"Cola" – not a brand. "Coke" – one syllable.

Wait, "Adidas" – three syllables. "Nike" – two syllables but four letters.

Wait, "Reebok" – six letters. "Fila" – four.

Wait, maybe a car brand. "Ford" – one syllable. "Chevy" – two syllables but five letters? C-H-E-V-Y. Five letters, "Chev" and "y"? Not opposites.

Wait, "Honda" – five letters, two syllables: Hon-da. Not opposites. "Toyota" – too long. "Lexus" – two syllables but five letters. "Le-xus". Not opposites.

Wait, "Nissan" – two syllables, six letters. "Mazda" – two syllables, five letters. M-A-Z-D-A. "Maz" and "da". Doesn't fit.

Wait, "Dodge" – one syllable. "Jeep" – one.

Wait, "Subaru" – three syllables.

Hmm. Let's think of gas stations again. "Shell", "Exxon", "Mobil", "Texaco", "Chevron", "BP", "Valero", "Citgo".

Citgo: C-I-T-G-O. Two syllables: Cit-go. "Cit" (like "sit") and "go". Wait, "sit" and "go" – those are opposites in a way. If you sit, you're not moving; if you go, you're moving. So "Citgo" – "Cit" phonetically is "sit", and "go" is spelled as "go". So the first syllable sounds like "sit", which is the opposite of "go". That might be it!

Citgo is a gas station brand, seen on highways, five letters (C-I-T-G-O), two syllables: Cit-go. Phonetically, "Cit" sounds like "sit", and "go" is the opposite. So the first syllable (phonetically "sit") is the opposite of the second syllable spelled "go". That seems to fit.

Let me double-check. The brand is Citgo. Five letters, two syllables. First syllable "Cit" pronounced like "sit", which is the opposite of "go". Yes, that makes sense. I think that's the answer.

The brand name is *Citgo*.

Citgo is a well-known gas station brand commonly seen along U.S. roads and highways. It has five letters and two syllables: *Cit* (pronounced like "sit") and *go*. The first syllable phonetically represents "sit," which is the opposite of the second syllable spelled as "go." This clever play on opposites fits the riddle's criteria perfectly.

Re: PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models

#16

[flagged]

I get that sliding in references to a passion project on top-scoring articles might seem like an easy way to give the project exposure, but commenting the same thing over and over comes off as a bit boorish. And just plugging the URL isn’t really contributing anything to the discussions IMO. Why not show us something your tool explained or summarized from the articles that isn’t obvious from a cursory read? Citing th…

I won't do it again my sincere apologies

Re: PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models

#17
As if the whole anti-intellectual hunt wasn’t enough, now PhD is a category implying holder of rote-knowledge at the highest level. I guess it is hopeless to fight this, but a PhD is 100x more about the apprenticeship and real-world training as a scientist than any accumulated knowledge beyond ones prior training.

I know this is a rant, sorry, just so tired of the stupidity.

Re: PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models

#18
post #15
post #13

My feeling is that a lot of challenge could come from the tokenizer used by the model, similar to r in strawberry problem.

There is likely a little bit of that, but I don't think that's the primary factor that makes these hard. If you look at the reasoning performed by R1 ("Show Thoughts" checkbox), it has no problem splitting and joining candidate words into letters, or even doing the letter-level manipulations. Just to give an example: Take a familiar brand name, seen along roads and highways in the United States. It has five letters,…

I see, but still there's a lot of reasonings just for counting the letters. And ridiculous reasonings like:

FedEx"? Five letters? F-E-D-E-X? No, that's six. Wait, FedEx is two syllables: Fed-Ex. Five letters? F-E-D-E-X? That's five letters? F, E, D, E, X. Yes, five letters. FedEx.

Re: PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models

#19
This doesn't feel like a "reasoning" challenge. The mental skill required to solve most of these seems to be the ability to loop over all known members of a category like "popular brand names" or "well-known actors" and see if they fit the clue.

As a human, you'd expect to fail either because you didn't know a category member (e.g. as a non-American I have no idea WTF "Citgo" is; I could never get the answer to the first question because I have never seen that name before in my life) or because you weren't able to bring it to mind; the mental act of looping over all members of a category is quite challenging for a human.

Admittedly this is something an AI system could in principle be REALLY good at, and it's interesting to test and see that current ones are not! But it seems weird to me to call what's being tested "reasoning" when it's so heavily focused on memory recall (and evaluating whether a candidate answer works or not is trivial once you've brought it to mind and doesn't really require any intelligent thought).

(If the questions were multiple-choice, eliminating the challenge of bringing candidate answers to mind that is the main challenge for a human, then I'd agree it was a "reasoning" test.)

Re: PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models

#20
post #18
post #15

Earlier quoted context omitted.

There is likely a little bit of that, but I don't think that's the primary factor that makes these hard. If you look at the reasoning performed by R1 ("Show Thoughts" checkbox), it has no problem splitting and joining candidate words into letters, or even doing the letter-level manipulations. Just to give an example: Take a familiar brand name, seen along roads and highways in the United States. It has five letters,…

I see, but still there's a lot of reasonings just for counting the letters. And ridiculous reasonings like: FedEx"? Five letters? F-E-D-E-X? No, that's six. Wait, FedEx is two syllables: Fed-Ex. Five letters? F-E-D-E-X? That's five letters? F, E, D, E, X. Yes, five letters. FedEx.

Definitely a lot of letter counting. It's not not a factor. I think the real problem is that the search space for each problem is enormous. When it gets stuck, it just gets stuck enumerating candidates that meet some but not all of the constraints.
Post reply on HN