Live data from Hacker News

Simple tasks showing reasoning breakdown in state-of-the-art LLMs

arxiv.org

391–393 of 393 posts

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#391
post #87

Earlier quoted context omitted.

I appear to be reasoning at times but I have mostly no idea what I am talking about. I hit a bunch of words and concepts in the given context and thus kind of hallucinate sense. Given a few months of peace of mind and enough money for good enough food, I could actually learn to reason without sounding like a confused babelarian. Reasoning is mostly a human convention supported by human context that would have been a…

Yeah, I think these chatbots are just too sure of themselves. They only really do "system 1 thinking" and only do "system 2 thinking" if you prompt them to. If I ask gpt-4o the riddle in this paper and tell it to assume its reasoning contains possible logical inconsistencies and to come up with reasons why that might be then it does correctly identify the problems with its initial answer and arrives at the correct on…

I'm not gonna read that book. I started and stopped after few chapters because it is based on and aims at manufacturing minds that follow game theory logic. Science (studies, reviews and application) got damaged quite a bit when too many people started following game theory logic.

We are, from our aware POV, a very young civilization.

And you only ever need game theory logic when you have to survive, got no thing and no skill to trade and you are too pathetic to move back in with your parents to work on your mind and or fuckability. Making money by ways of game theory logic compensates for all that but also diminishes the survival chance of the users' offspring to zero once super-unalligned AGIs start to assess the entire supply chain of wealth and how it impacts the evolution of human organisms and the ones inside them.

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#392

Earlier quoted context omitted.

Yeah, I think these chatbots are just too sure of themselves. They only really do "system 1 thinking" and only do "system 2 thinking" if you prompt them to. If I ask gpt-4o the riddle in this paper and tell it to assume its reasoning contains possible logical inconsistencies and to come up with reasons why that might be then it does correctly identify the problems with its initial answer and arrives at the correct on…

If you had a prompt that reliably made the model perform better at all tasks, that would be useful. But if you have to manually tweak your prompts for every problem, and then manually verify that the answer is correct, that's not so useful.

the fact that you can manually tweak your prompts for any problem and agent, is still super useful (joke: *and the only reason our civilization still exists)

Re: Simple tasks showing reasoning breakdown in state-of-the-art LLMs

#393

Earlier quoted context omitted.

Yeah, I think these chatbots are just too sure of themselves. They only really do "system 1 thinking" and only do "system 2 thinking" if you prompt them to. If I ask gpt-4o the riddle in this paper and tell it to assume its reasoning contains possible logical inconsistencies and to come up with reasons why that might be then it does correctly identify the problems with its initial answer and arrives at the correct on…

There isn't any evidence that models are doing any kind of "system 2 thinking" here. The model's response is guided by both the prompt and its current output so when you tell it to reason step by step the final answer is guided by its current output text. The second best answer is just something it came up with because you asked, the model has no second best answer to give. The second best answers always seem strange…

well, the tuning of training data results in at least some predictions that resemble varying models of systems 1 and 2 thinking. there is no reasoning at all. it's all models of reasoning, tokenized by opinionated taxonomical algorithms and degrees of systemic, academic/conventional human interpretation (tags) that are far from capturing the general human experience.
Post reply on HN