Live data from Hacker News

I got the highest score on ARC-AGI again swapping Python for English

jeremyberman.substack.com

111–120 of 136 posts

Re: I got the highest score on ARC-AGI again swapping Python for English

#111

Earlier quoted context omitted.

> are they really learning to reason, or are they just learning to pattern match to steer generation in the direction of problem-specific reasoning steps that they had been trained on? Are you sure there's a real difference? Do you have a definition of "reasoning" that excludes this?

I define intelligence as prediction (degree of ability to use past experience to correctly predict future action outcomes), and reasoning/planning as multi-step what-if prediction. Certainly if a human (or some AI) has learned to predict/reason over some domain, then what they will be doing is pattern matching to determine the generalizations and exceptions that apply in a given context (including a hypothetical cont…

> but rather the ability to reason in the general case, which requires the ability to LEARN to solve novel problems, which is what is missing from LLMs.

I don't think it's missing, zero shot prompting is quite successful in many cases. Maybe you find the extent that LLMs can do this to be too limited, but I'm not sure that means they don't reason at all.

> A system that has a fixed set of (reasoning/prediction) rules, but can't learn new ones for itself, seems better regarded as an expert system.

I think expert systems are a lot more limited than LLMs, so I don't agree with that classification. LLMs can generate output that's out of distribution, for instance, which is not something that's classic expert systems can do (even if you think LLM OOD is still limited compared to humans).

I've elaborated in another comment [1] what I think part of the real issue is, and why people keep getting tripped up by saying that pattern matching is not reasoning. I think it's perfectly fine to say that pattern matching is reasoning, but pattern matching has levels of expressive power. First-order pattern matching is limited (and so reasoning is limited), and clearly humans are capable of higher order pattern matching which is Turing complete. Transformers are also Turing complete, and neural networks can learn any function, so it's not a matter of expressive power, in principle.

Aside from issues stemming from tokenization, I think many of these LLM failures are because they aren't trained in higher order pattern matching. Thinking models and the generalization seen from grokking are the first steps on this path, but it's not quite there yet.

[1] https://news.ycombinator.com/item?id=45277098

Re: I got the highest score on ARC-AGI again swapping Python for English

#112
post #87

Earlier quoted context omitted.

> It's trivial to demonstrate that LLMs are pattern matching rather than reasoning. Again, this is just asserting the premise that reasoning cannot include pattern matching, but this has never been justified. What is your definition for "reasoning"? > This is clearly pattern matching and overfitting to the "doctor riddle" and a good demonstration of how there's no actual reasoning going on. Not really, no. "Bad reaso…

If your assertion is that you can't prove reasoning isn't just pattern matching, then I counter by saying you can't prove reasoning isn't just chaining a large number of IF/THEN/ELSE logic statements and therefore computers have been generally intelligent since ~1960.

The difference between ML models and computers since the 1960s is that the ML models weren't programmed with predicates, they "learned" them from analyzing data, and can continue to learn in various ways from further data. That's a meaningful difference, and why the former may qualify as intelligent and the latter cannot.

But I agree in principle that LLMs can be distilled into large IF/THEN/ELSE trees, that's the lesson of BitNet 1-bit LLMs. The predicate tree being learned from data is the important qualifier for intelligence though.

Edit: in case I wasn't clear, I agree that a specific chain of IF/THEN/ELSE statements in a loop can be generally intelligent. How could it not, specific kinds of these chains are Turing complete after all, so unless you think the brain has some kind of magic, it too is reducible to such a program, in principle. We just haven't yet discovered what kind of chain this is, just like we didn't understand what kind of chain could produce distributed consensus before PAXOS.

Re: I got the highest score on ARC-AGI again swapping Python for English

#113

Earlier quoted context omitted.

Do you?

https://en.wikipedia.org/wiki/Commonsense_reasoning

So, you don't, but wikipedia does? I'll believe they can do commonsense reasoning when they can figure out that people have 4 fingers and 1 thumb. Here I was thinking common sense reasoning was what we call reasoning based on common sense. Go figure some AI folks needed to write a wikipedia article to redefine common sense.

Like they say, common sense ain't so common at all.

Re: I got the highest score on ARC-AGI again swapping Python for English

#114

Earlier quoted context omitted.

https://en.wikipedia.org/wiki/Commonsense_reasoning

So, you don't, but wikipedia does? I'll believe they can do commonsense reasoning when they can figure out that people have 4 fingers and 1 thumb. Here I was thinking common sense reasoning was what we call reasoning based on common sense. Go figure some AI folks needed to write a wikipedia article to redefine common sense. Like they say, common sense ain't so common at all.

Least you could do is look up what an unfamiliar term means before rolling in with all the hot takes.

So take the link, and read it. That would help you to be less ignorant the next time around.

Re: I got the highest score on ARC-AGI again swapping Python for English

#115

Earlier quoted context omitted.

It seems readily apparent there is a difference given their inability to do tasks we would otherwise reasonably describe as achievable via basic reasoning on the same facts.

I agree LLMs have many differences in abilities relative to humans. I'm not sure what this implies for their ability to reason though. I'm not even sure what examples about their bad reasoning can prove about the presence or absence of any kind of "reasoning", which is why I keep asking for definitions to remove the ambiguity. If examples of bad reasoning sufficed, then this would prove that humans can't reason eithe…

Nothing about the uncertainty of the definition for 'reasoning' requires that pattern matching be part of the definition.

Re: I got the highest score on ARC-AGI again swapping Python for English

#116

Earlier quoted context omitted.

So, you don't, but wikipedia does? I'll believe they can do commonsense reasoning when they can figure out that people have 4 fingers and 1 thumb. Here I was thinking common sense reasoning was what we call reasoning based on common sense. Go figure some AI folks needed to write a wikipedia article to redefine common sense. Like they say, common sense ain't so common at all.

Least you could do is look up what an unfamiliar term means before rolling in with all the hot takes. So take the link, and read it. That would help you to be less ignorant the next time around.

>Least you could do is look up what an unfamiliar term means before rolling in with all the hot takes.

Thanks for proving my point that common sense ain't so common. To be clear, common sense reasoning is not an "unfamiliar term" save for this new (article was written in 2021) redefinition of it to be something AI related. It's kinda laughable that you are being this snitty about.

> That would help you to be less ignorant the next time around.

Better to be "ignorant" than slow and humorless.

Re: I got the highest score on ARC-AGI again swapping Python for English

#117
post #68

Earlier quoted context omitted.

> are they really learning to reason, or are they just learning to pattern match to steer generation in the direction of problem-specific reasoning steps that they had been trained on? Are you sure there's a real difference? Do you have a definition of "reasoning" that excludes this?

It's trivial to demonstrate that LLMs are pattern matching rather than reasoning. A good way is to provide modified riddles-that-aren't. As an example: > Prompt: A man working at some white collar job gets an interview scheduled with an MBA candidate. The man says "I can't interview this candidate, he's my son." How is this possible? > ChatGPT: Because the interviewer is the candidate’s mother. (The riddle plays on t…

OK. But, in Claude Sonnet 4:

'This is possible because the man is the candidate's father. When he says "he's my son," he's simply stating their family relationship. The scenario doesn't present any logical contradiction - a father could very well be in a position where he's supposed to interview his own son for a job. This would create a conflict of interest, which is why he's saying he can't conduct the interview. It would be inappropriate and unfair for a parent to interview their own child for a position, so he would need to recuse himself and have someone else handle the interview. The phrasing might initially seem like it's setting up a riddle, but it's actually a straightforward situation about professional ethics and avoiding conflicts of interest in hiring.'

EDIT - this is described better by other posters.

Re: I got the highest score on ARC-AGI again swapping Python for English

#118

Earlier quoted context omitted.

What's obtuse about it? It's honestly a very straightforward statement. Every thing we think or say is a function of past events. We don't incorporate future events into what we think or say. Even speculation or imagination of future events occurred in the past (that is the act of imagining it occurred in the past). It's really a super simple concept -- maybe it's so simple that it seems obtuse.

Because the other poster's point wasn't that it was a 'past event.' The point was that it's just predicting based upon the previous token . It's disingenuous to mix the two concepts up.

> The point was that it's just predicting based upon the previous token.

Well that's just wrong. None of the LLMs of interest predict based upon the previous token.

Re: I got the highest score on ARC-AGI again swapping Python for English

#119

Earlier quoted context omitted.

I agree LLMs have many differences in abilities relative to humans. I'm not sure what this implies for their ability to reason though. I'm not even sure what examples about their bad reasoning can prove about the presence or absence of any kind of "reasoning", which is why I keep asking for definitions to remove the ambiguity. If examples of bad reasoning sufficed, then this would prove that humans can't reason eithe…

Nothing about the uncertainty of the definition for 'reasoning' requires that pattern matching be part of the definition.

Did someone in this thread claim that?

Re: I got the highest score on ARC-AGI again swapping Python for English

#120

Earlier quoted context omitted.

I define intelligence as prediction (degree of ability to use past experience to correctly predict future action outcomes), and reasoning/planning as multi-step what-if prediction. Certainly if a human (or some AI) has learned to predict/reason over some domain, then what they will be doing is pattern matching to determine the generalizations and exceptions that apply in a given context (including a hypothetical cont…

> but rather the ability to reason in the general case, which requires the ability to LEARN to solve novel problems, which is what is missing from LLMs. I don't think it's missing, zero shot prompting is quite successful in many cases. Maybe you find the extent that LLMs can do this to be too limited, but I'm not sure that means they don't reason at all. > A system that has a fixed set of (reasoning/prediction) rules…

Powerful pattern matching is still just pattern matching.

How is an LLM going to solve a novel problem with just pattern matching?

Novel means it has never seen it before, maybe doesn't even have the knowledge needed to solve it, so it's not going to be matching any pattern, and even if it did, that would not help if it required a solution different to whatever the pattern match had come from.

Human level reasoning includes ability to learn, so that people can solve novel problems, overcome failures by trial and error, exploration, etc.

So, whatever you are calling "reasoning" isn't human level reasoning, and it's therefore not even clear what you are trying to say? Maybe just that you feel LLMs have room for improvement by better pattern matching?

Post reply on HN