Live data from Hacker News

The case for zero-error horizons in trustworthy LLMs

arxiv.org

21–30 of 121 posts

Re: The case for zero-error horizons in trustworthy LLMs

#21
post #20

Doesn't this just look like another case of "count the r's in strawberry" ie not understanding how tokenization works? This is well known and not that interesting to me - ask the model to use python to solve any of these questions and it will get it right every time.

It's not dismissible as a misunderstanding of tokens. LLMs also embed knowledge of spelling - that's how they fixed the strawberry issue. It's a valid criticism and evaluation.

Re: The case for zero-error horizons in trustworthy LLMs

#22
To those saying this is not surprising, yes it will be surprising to the general public who are being served ads from huge companies like MS or OpenAI saying LLMs can help with their accounting, help them close deals by crunching the numbers in seconds, write complex code for them etc etc.

This is important information for anyone to understand who thinks these systems are thinking, reasoning, and learning from them or that they’re having a conversation with them i.e. 90% of users of LLMs.

Re: The case for zero-error horizons in trustworthy LLMs

#23

People are going to misinterpret this and overgeneralize the claim. This does not say that AI isn't reliable for things. It provides a method for quantifying the reliability for specific tasks. You wouldn't say that a human who doesn't know how to read isn't reliable in everything, just in reading. Counting is something that even humans need to learn how to do. Toddlers also don't understand quantity. If a 2 year old…

You’re conflating counting and language.

Many animals can count. Counting is recognizing that the box with 3 apples is preferable to the one with 2 apples.

Yes, 2 year olds might struggle with the externalization of numeric identities but if you have 1 M&M in one hand and 5 in the other and ask which they want, they’ll take the 5.

LLMs have the language part down, but fundamentally can’t count.

Re: The case for zero-error horizons in trustworthy LLMs

#24
post #16

Whenveer I see these papers and try them, they always work. This paper is two months old, which in LLM years is like 10 years of progress. It would be interesting to actively track how far long each progressive model gets...

Even more interesting to track how many of those are just ad-hoc patched.

Probably zero. At the end of the day people pay for LLMs that write better code or summarize PDFs of hundreds of pages faster, not the ones that can count the letter r's better.

When LLMs can't count r's: see? LLMs can't think. Hoax!

When LLMs count r's: see? They patched and benchmark-maxxed. Hoax!

You just can't reason with the anti-LLM group.

Re: The case for zero-error horizons in trustworthy LLMs

#26
post #23

People are going to misinterpret this and overgeneralize the claim. This does not say that AI isn't reliable for things. It provides a method for quantifying the reliability for specific tasks. You wouldn't say that a human who doesn't know how to read isn't reliable in everything, just in reading. Counting is something that even humans need to learn how to do. Toddlers also don't understand quantity. If a 2 year old…

You’re conflating counting and language. Many animals can count. Counting is recognizing that the box with 3 apples is preferable to the one with 2 apples. Yes, 2 year olds might struggle with the externalization of numeric identities but if you have 1 M&M in one hand and 5 in the other and ask which they want, they’ll take the 5. LLMs have the language part down, but fundamentally can’t count.

The concept of bigger/smaller is useful but is a distinct skill from counting. If you spread the M&Ms apart enough that the part of the brain responsible for gestalt clustering can't group them into a "bigger whole" signal, they'll no longer be able to do the thing you're saying (this is the law of proximity in gestalt psychology).

Re: The case for zero-error horizons in trustworthy LLMs

#29

There’s no way this is right. I checked complicated ones with the latest thinking model. Can someone come up with a counter example? Edit: here’s what I tried https://chatgpt.com/share/69cebb52-56a8-838f-969c-c47308262a...

"in this paper we primarily evaluate the LLM itself without external tool calls."

Maybe this is a factor?

Post reply on HN