Live data from Hacker News

Simple tasks showing reasoning breakdown in state-of-the-art LLMs

arxiv.org

381–390 of 393 posts

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#381

Earlier quoted context omitted.

Most people can't do 1 + 3x = 22 without any words or symbols. People who can don't realize that most people can't. I'd argue one isn't using logic when they do that, it's just very good pattern matching.

It's also possible to do mentally by visualizing it rather than internal monologue. You can imagine the 1 on the left arcing over to the right, cancelling the 22 to 21, then the 3 moving under the 21 and the 21 descending through the 3 to become 7.

yup. I considered myself an /extremely/ verbal person when reasoning, but what I do with the above feels closest to 'moving the 1', almost like balancing a mental scale.

I never really noticed that before. I'm not great at math, fwiw.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#382
post #381

Earlier quoted context omitted.

It's also possible to do mentally by visualizing it rather than internal monologue. You can imagine the 1 on the left arcing over to the right, cancelling the 22 to 21, then the 3 moving under the 21 and the 21 descending through the 3 to become 7.

yup. I considered myself an /extremely/ verbal person when reasoning, but what I do with the above feels closest to 'moving the 1', almost like balancing a mental scale. I never really noticed that before. I'm not great at math, fwiw.

A decent number of folks just have the answer pop into their head, no logic or thinking required.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#383

Citation 40 is the longest list of authors I have ever seen. That is one way to help all your friends get tenure.

Wow! You weren't kidding. People in my field often joke that papers in biology, health etc. tend to list every person they met on the day of submission as authors. But this one has 450 authors from 132 institutions. Surely this points to some sort of breakdown in the attribution system.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#384

Earlier quoted context omitted.

Do you have any concern about the data you're feeding to the vendor serving your prompts? I've had junior devs tell me they use chatgippity to combine excel workbooks, and when I confirm they're not self hosting a llm to do it, I ask if they think it's a good idea to hand over company data to openai. They don't care. In a world of tight security, I find it astonishing that so many people willingly give away trade sec…

So you are not using Office 365? Because our company does.

Has something changed with the service agreement? I was under the understanding that Microsoft didn't mine or sell to advertisers 365 data, at least for corporate accounts.

https://techcommunity.microsoft.com/t5/microsoft-365-blog/wh...

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#385

Earlier quoted context omitted.

That, as I understand it, is not a valid chain of logic. Requiring fewer data points does not inherently indicate that the underlying mechanism (autogressive sequential generation, not the transformer which is just an architecture) is different. Not to mention the secondary arguments like - no proof that human learns faster from fewer datapoints, that's just your assumption in the sibling comment. Humans inherit info…

Well, the same thing goes for you - just because someone posts on HN doesn't mean they know what they're talking about. And if I have to decide whose assessment I trust regarding AI, I take the Turing award winner who worked for almost 40 years on AI over a random guy from the internet.

Sure. But I'm not standing here saying my argument is valid because SolidAsparagus made it.

> And if I have to decide whose assessment I trust regarding AI, I take the Turing award winner who worked for almost 40 years on AI over a random guy from the internet.

I'd encourage you to do your own thinking and make up your own mind.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#386

Earlier quoted context omitted.

AI needs to see thousands or millions of images of a cat before they reliably can identify one. The fact that a child needs to only see one example of a cat to know what a cat is from then on seems to point to humans having something very different.

Humans train on continuous video . Even our most expensive models are, in terms of training set size, far behind what an infant processes in the first year of their life. EDIT: and it takes human children a couple years to reliably identify a cat. My 2.5 y.o. daughter still confuses cats with small dogs, despite living under one roof with a cat.

It's a massive amount of data. Not just video, but all senses, continuously. Also, I think an important aspect is that the training data is interactive. From day one, an infant sees his own hands moving and begins testing cause and effect. They aren't passively absorbing.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#387

Earlier quoted context omitted.

So you are not using Office 365? Because our company does.

Has something changed with the service agreement? I was under the understanding that Microsoft didn't mine or sell to advertisers 365 data, at least for corporate accounts. https://techcommunity.microsoft.com/t5/microsoft-365-blog/wh...

But thats the point. You can use OpenAI through Azure and other means. If you trust MS (who has basically data from millions of companies) why would it not work out with OpenAI usage?

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#388
So interestingly if you are more explicit all of the models reliably answer correct. The confusion is clearly whether Alice is included as a sister in the count of sisters. By stipulating the prompt with “not including herself” it sufficiently disambiguates the problem and most LLMs get it right.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#389
post #153

Earlier quoted context omitted.

> People are being told it can do far more than it really is. Meanwhile these HN comments are split between: * Lots of people confirming what the paper itself notes (but doesn't highlight), that the most advanced models actually can solve this problem at least a significant portion of the time. (A proportion which one can pretty easily project is only likely to increase with future models.) * Lots of people saying "t…

The critical understanding doesnt predict that LLMs cannot solve problems. It predicts how they will solve them. There is no information, a priori, what the LLM has been trained on. You have to prompt, then see the answer. Once the answer arrives, the critical understanding provides a route to repairing the answer when not accurate or useful. LLMs do not reason. They appear to reason by repeating the structure of rea…

I am very curious how you would define 'reasoning'.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#390

Earlier quoted context omitted.

Yeah, I think these chatbots are just too sure of themselves. They only really do "system 1 thinking" and only do "system 2 thinking" if you prompt them to. If I ask gpt-4o the riddle in this paper and tell it to assume its reasoning contains possible logical inconsistencies and to come up with reasons why that might be then it does correctly identify the problems with its initial answer and arrives at the correct on…

> After you answer the riddle please review your answer assuming that you have made a logical inconsistency in each step and explain what that inconsistency is. Even if you think there is none do your best to confabulate a reason why it could be logically inconsistent. LLMs are fundamentally incapable of following this instruction. It is still model inference, no matter how you prompt it.

isn't any instruction a subclass of inference? and doesn't any phrasing (lexicology) simply translate "down" to the heaviest values which, varying with the fine tuning, are the words that are, consensually and conventionally, the simplest ones that convey the meaning of the original word in the prompt, which should be the least ambivalent/least interpretable (again, fine tuning can broaden the scope) oneS. thus the LLM fulfills the "translated" instructions step by step and comes up with both or either the correct reasoning and answer.

details and technicalities, especially liminal ones, aren't as conventional and consensual as the name of the current set it is to be interpreted in.

so almost all mistakes of LLMs can be blamed on the lack of variety of human translations. multiple translations are only common for subtitles, mangas and manhwa as far as i know, or when some dude or dudette is proficient and passionate in two languages and reads a bad/weak translation of a (usually classic) novel. why the fuck would a human properly retranslate automated documentations or googles dev blog? or books on logic, in any science, books on art and aesthetics and whatnot. technical people don't need to care because, practically, there are no interpretations in algorithms and the rest of the code, except when a programming language does something weird on the (or someones) machine, which isn't that common by design.

Post reply on HN