Earlier quoted context omitted.
People aren't claiming that they are holding textbooks in their model, that would just be even more evidence of reasoning (the LLM would have to reason what textbook to reference, and then extrapolate from the textbook(s) how to solve the problem at hand - pretty much what students in school do; study the textbook and reason from it to answer new test questions) People are claiming that the models sit on a vast archi…
I think the more realistic argument is that the model can generalize, but only by learning shortcuts (e.g. how to pattern match a problem to a likely answer) and simple algorithms (e.g. how to propagate carries in a multiplication). And this only appears intelligent because this pattern matching is really good and backed by a huge amount of compressed/memorized answers. The AI pessimist's argument is that there's a h…
Seven replies to the viral Apple reasoning paper and why they fall short
301–310 of 331 posts
Re: Seven replies to the viral Apple reasoning paper and why they fall short
#302> Puzzles a child can do Certainly, I couldn't solve Hanoi's towers with 8 disks purely in my mind without being able to write down the state of every step or having a physical state in front of me. Are we comparing apples to apples?
Re: Seven replies to the viral Apple reasoning paper and why they fall short
#303As we are losing our rights in America(won't even acknowledge the 'new knowledge'), this becomes important to freedom loving people of the world. Please, acknowledge this important work for the world. This 'new knowledge' is free to the world. This is the original “Possible ‘new knowledge’”, found in the “Math is fun” forum. All files can be found at: https://drive.google.com/drive/folders/1wpd5-2-4SZkZka284sbp... Ma…
Re: Seven replies to the viral Apple reasoning paper and why they fall short
#304Earlier quoted context omitted.
>with tool use A LLM with tool use can solve anything. It is interesting to try and measure its capabilities without tools.
I don't think the first is true at all, unless you imagine some powerful oracle tools. I think the second is interesting for comparing models, but not interesting for determining the limits of what models can automate in practice. It's the prospect of automating labour which makes AI exciting and revolutionary, not their ability when arbitrarily restricted.
It would draw on many previously written examples of algorithms to write the code for solving Hanoi. To solve a novel problem with tool use, one needs to work sequentially while staying on task, notice where you've gone wrong, and backtrack.
I don't want to overstate the case here, I'm sure there is work where there's enough intersection between previously existing stuff in the dataset and few enough sequential steps required that useful work can be done, but idk how much you've tried using this stuff as a labour saving device, there's less low hanging fruit than one might think, but more than zero.
Re: Seven replies to the viral Apple reasoning paper and why they fall short
#305> 1. Humans have trouble with complex problems and memory demands. True! But incomplete. We have every right to expect machines to do things we can’t. [...] If we want to get to AGI, we will have to better. I don't get this argument. The paper is about "whether RLLMs can think". If we grant "humans make these mistakes too", but also "we still require this ability in our definition of thinking", aren't we saying "thin…
Re: Seven replies to the viral Apple reasoning paper and why they fall short
#306Earlier quoted context omitted.
> That is wishful thinking popularised by Ilya Sutskever Ilya and Hinton have claimed even crazier things | to understand next token prediction you must understand the casual reality This is objectively false. It's a result known in physics to be wrong for centuries. You can probably reason a weaker case yourself, that I'm sure you can make accurate predictions about some things without fully understanding them. But…
Hinton and Sutskever are victims of their own success: they can say whatever they like and nobody dares criticise them, or tell them how they're wrong. I recently watched a video of Sutskever speaking to some students, not sure where and I can't dig out the link now. To summarise he told them that the human brain is a biological computer. He repeated this a couple of times then said that this is why we can create a d…
Re: Seven replies to the viral Apple reasoning paper and why they fall short
#307Earlier quoted context omitted.
I'm replying to this > the model printing out a bunch of steps, saying there is no point in doing it thousands of times more.
Ok, but the very next sentence was: > And they they'd either output an the algorithm for printing the rest in words or code. So clearly you already knew that your strawman was not relevant. Why try it anyway?
Re: Seven replies to the viral Apple reasoning paper and why they fall short
#308How can we then assess if machine is doing it?
As Demis Hassabis put it a while back: we are building AI [partially] to understand how our own brain works.
Re: Seven replies to the viral Apple reasoning paper and why they fall short
#309Earlier quoted context omitted.
Yes, but “good at” here has a very limited, technical meaning, which can be oversimplified as “better than random chance.” If something can be better than random chance in any arbitrary problem domain it was not trained on, that is AGI.
That raises the plausible question; are there problem domains where humans cannot do better than random chance, given repeated attempts?
Re: Seven replies to the viral Apple reasoning paper and why they fall short
#310Earlier quoted context omitted.
It might not fit your work, but there are tons of areas where “good enough” can still provide a lot of value. I’m sure you’d be thrilled with a tool that could correctly tell you if Apple’s stock was going up or down tomorrow 70% of the time.
I work in a mail room sending hard copy letters to customers. If I got my job right only 70% of the time then I’d be causing massive privacy breaches daily by sending the wrong personal information to the wrong customers. Would you trust an AI that gets your banking transactions right only 70% of the time?