Doesn't this just look like another case of "count the r's in strawberry" ie not understanding how tokenization works? This is well known and not that interesting to me - ask the model to use python to solve any of these questions and it will get it right every time.
The case for zero-error horizons in trustworthy LLMs
21–30 of 121 posts
Re: The case for zero-error horizons in trustworthy LLMs
#22This is important information for anyone to understand who thinks these systems are thinking, reasoning, and learning from them or that they’re having a conversation with them i.e. 90% of users of LLMs.
Re: The case for zero-error horizons in trustworthy LLMs
#23People are going to misinterpret this and overgeneralize the claim. This does not say that AI isn't reliable for things. It provides a method for quantifying the reliability for specific tasks. You wouldn't say that a human who doesn't know how to read isn't reliable in everything, just in reading. Counting is something that even humans need to learn how to do. Toddlers also don't understand quantity. If a 2 year old…
Many animals can count. Counting is recognizing that the box with 3 apples is preferable to the one with 2 apples.
Yes, 2 year olds might struggle with the externalization of numeric identities but if you have 1 M&M in one hand and 5 in the other and ask which they want, they’ll take the 5.
LLMs have the language part down, but fundamentally can’t count.
Re: The case for zero-error horizons in trustworthy LLMs
#24Whenveer I see these papers and try them, they always work. This paper is two months old, which in LLM years is like 10 years of progress. It would be interesting to actively track how far long each progressive model gets...
Even more interesting to track how many of those are just ad-hoc patched.
When LLMs can't count r's: see? LLMs can't think. Hoax!
When LLMs count r's: see? They patched and benchmark-maxxed. Hoax!
You just can't reason with the anti-LLM group.
Re: The case for zero-error horizons in trustworthy LLMs
#25Re: The case for zero-error horizons in trustworthy LLMs
#26People are going to misinterpret this and overgeneralize the claim. This does not say that AI isn't reliable for things. It provides a method for quantifying the reliability for specific tasks. You wouldn't say that a human who doesn't know how to read isn't reliable in everything, just in reading. Counting is something that even humans need to learn how to do. Toddlers also don't understand quantity. If a 2 year old…
You’re conflating counting and language. Many animals can count. Counting is recognizing that the box with 3 apples is preferable to the one with 2 apples. Yes, 2 year olds might struggle with the externalization of numeric identities but if you have 1 M&M in one hand and 5 in the other and ask which they want, they’ll take the 5. LLMs have the language part down, but fundamentally can’t count.
Re: The case for zero-error horizons in trustworthy LLMs
#27Edit: here’s what I tried https://chatgpt.com/share/69cebb52-56a8-838f-969c-c47308262a...
Re: The case for zero-error horizons in trustworthy LLMs
#28Re: The case for zero-error horizons in trustworthy LLMs
#29There’s no way this is right. I checked complicated ones with the latest thinking model. Can someone come up with a counter example? Edit: here’s what I tried https://chatgpt.com/share/69cebb52-56a8-838f-969c-c47308262a...
Maybe this is a factor?
Re: The case for zero-error horizons in trustworthy LLMs
#30[flagged]