Live data from Hacker News

Apple study proves LLM-based AI models are flawed because they cannot reason

appleinsider.com

11–20 of 24 posts

Re: Apple study proves LLM-based AI models are flawed because they cannot reason

#12
post #5

Gosh I wish someone would pay me handsomely for coming up with such stupidly obvious "research" results as "a computer program that uses statistics to pick the next word in a sequence doesn't reason like a person".

VCs spend obscene amounts of money trying to manufacture consumer demand for a fake product, then we are left arguing for years about it. Such a waste of time.

Re: Apple study proves LLM-based AI models are flawed because they cannot reason

#13
post #5

Gosh I wish someone would pay me handsomely for coming up with such stupidly obvious "research" results as "a computer program that uses statistics to pick the next word in a sequence doesn't reason like a person".

There’s clearly enough people who need to be told this, given any thread on this topic in recent years

Re: Apple study proves LLM-based AI models are flawed because they cannot reason

#14
post #12
post #5

Gosh I wish someone would pay me handsomely for coming up with such stupidly obvious "research" results as "a computer program that uses statistics to pick the next word in a sequence doesn't reason like a person".

VCs spend obscene amounts of money trying to manufacture consumer demand for a fake product, then we are left arguing for years about it. Such a waste of time.

ChatGPT got hundreds of millions of users with no advertising, and unlike (say) Google+, OpenAI didn't have a pre-existing userbase to push new products on. So there's clearly some demand.

Re: Apple study proves LLM-based AI models are flawed because they cannot reason

#15
post #4

This is silly. Humans will also get fewer right answers if you make the question more complex (requiring additional steps), or if you add irrelevant information as a distraction (since on tests, there's usually an assumption that all information given is relevant). As for changing names and numbers, the large effects they saw were all on small (<10B param) open source models; the effects on o1 were tiny and barely di…

>>> since on tests, there's usually an assumption that all information given is relevant

Maybe on grade-school tests, but the professional certification exams I've taken have questions about scenarios where part of the challenge is recognizing which parts of the scenario description aren't relevant. The one I had an in-person class for, the instructor specifically called this out and advised reading the questions first so we'd know which parts of the scenario we didn't have to think about while reading it.

Re: Apple study proves LLM-based AI models are flawed because they cannot reason

#16
post #5

Gosh I wish someone would pay me handsomely for coming up with such stupidly obvious "research" results as "a computer program that uses statistics to pick the next word in a sequence doesn't reason like a person".

You can do it too, just take a PhD in any field.

Re: Apple study proves LLM-based AI models are flawed because they cannot reason

#17

Interesting. I would have thought that the training set (basically the whole internet AIUI) would have included various "teacher's version" exams with enough word problems with intentionally-distracting extra information, that the models would be able to ignore that sort of thing. This sounds like they're inspecting existing models. Maybe a model trained specifically on "word problem" question-answer pairs (as in, th…

They hid away the results of o1 preview in the Appendix but it does not drop below margin of error numbers on 4/5 of their modified benchmarks. The last one they add "seemingly relevant but ultimately irrelevant information to problems" and it drops to 77%. Now I'm willing to bet this is within human baselines but either way, researchers really need to start including human baselines in these kinds of papers.

>>> I'm willing to bet this is within human baselines but either way, researchers really need to start including human baselines in these kinds of papers.

They should indeed, but is that "I've never seen this before" human baseline, or with prior exposure ("what do you mean I got that wro... oh I see what you did there") or explicit instruction?

Re: Apple study proves LLM-based AI models are flawed because they cannot reason

#19

Earlier quoted context omitted.

They hid away the results of o1 preview in the Appendix but it does not drop below margin of error numbers on 4/5 of their modified benchmarks. The last one they add "seemingly relevant but ultimately irrelevant information to problems" and it drops to 77%. Now I'm willing to bet this is within human baselines but either way, researchers really need to start including human baselines in these kinds of papers.

>>> I'm willing to bet this is within human baselines but either way, researchers really need to start including human baselines in these kinds of papers. They should indeed, but is that "I've never seen this before" human baseline, or with prior exposure ("what do you mean I got that wro... oh I see what you did there") or explicit instruction?

I mean baselines to match how the LLMs are being tested in whatever paper (as best as possible). In this case, average scores on the unaltered benchmark and then average scores on each modified benchmark to indicate how much on average human performance drops introducing these details.

A "hey, you got that wrong. check again" is fine if the LLMs in the paper are also being prompted that way.

Re: Apple study proves LLM-based AI models are flawed because they cannot reason

#20
post #5

Gosh I wish someone would pay me handsomely for coming up with such stupidly obvious "research" results as "a computer program that uses statistics to pick the next word in a sequence doesn't reason like a person".

AlphaProof was able to get a silver IMO medal by writing formally verified proofs of novel mathematical problems: https://deepmind.google/discover/blog/ai-solves-imo-problems... Whether this is "like a person" or not, it seems silly to insist that this "doesn't count" as mathematical reasoning. I certainly couldn't get an IMO silver medal and I have a degree in math.

Hang on a second, there's a lot wrong with this.

> First, the problems were manually translated into formal mathematical language for our systems to understand.

Ok so the "AI" wasn't solving the same problem as every other Olympiad, it was solving a "translated" version. I wonder how much of the solving was performed in this translation. If the model is so capable of reasoning, why was this step performed by people?

> In the official competition, students submit answers in two sessions of 4.5 hours each. Our systems solved one problem within minutes and took up to three days to solve the others.

So no, they would not have been awarded a silver medal if they were competing in the IMO.

Not to mention, AlphaProof is not an LLM and has absolutely nothing to do with what I was commenting about.

Post reply on HN