Live data from Hacker News

The case for zero-error horizons in trustworthy LLMs

arxiv.org

51–60 of 121 posts

Re: The case for zero-error horizons in trustworthy LLMs

#52

Earlier quoted context omitted.

I just tried it in ChatGPT "Auto" and it didn't work > Yes — ((((()))))) is balanced. > It has 6 opening ( and 6 closing ), and they’re properly nested. Though it did work when using "Extensive Thinking". The model wrote a Python program to solve this. > Almost balanced — ((((()))))) has 5 opening parentheses and 6 closing parentheses, so it has one extra ). > A balanced version would be: ((((())))) Testing a couple…

Weird. I tried in chatGPT auto and it worked perfectly. I tried like 10 variations. I also did the letters in words. Got all of them right. The one thing I did trip it up on was "Is there the sh sound in the word transportation". It said no. And then realized I asked for "sound" not letters. It then subsequently got the rest of the "sounds-like" tests I did. Clearly, my ChatGPT is just better than yours.

heh, interesting that. I just tried it twice more with ChatGPT "Instant" (disabling "Auto-switch to Thinking") and it got it wrong both times. Does yours get it right without thinking or tool calls? If so, maybe it does like you better than me.

Re: The case for zero-error horizons in trustworthy LLMs

#53
post #29

There’s no way this is right. I checked complicated ones with the latest thinking model. Can someone come up with a counter example? Edit: here’s what I tried https://chatgpt.com/share/69cebb52-56a8-838f-969c-c47308262a...

"in this paper we primarily evaluate the LLM itself without external tool calls." Maybe this is a factor?

No tools were used.

Re: The case for zero-error horizons in trustworthy LLMs

#54

To those saying this is not surprising, yes it will be surprising to the general public who are being served ads from huge companies like MS or OpenAI saying LLMs can help with their accounting, help them close deals by crunching the numbers in seconds, write complex code for them etc etc. This is important information for anyone to understand who thinks these systems are thinking, reasoning, and learning from them o…

Quick sanity check: you're susceptible to pretty irresistible optical illusions which would never fool a VLM, does it mean you're not thinking? In fact, with a non-monospaced font I also have trouble determining whether these parens are balanced, and have to select them with the mouse, i.e. use a "dumb" tool, to make sure.

Reminder that "thinking" is an ill-defined term like others, and the question whether they "think" is basically irrelevant. No intelligent system, human or machine, will ever have zero error rate, due to the very nature of intelligence (another vague term). You have to deal with that the same way you deal with it in humans - either treat bugs as bugs and build systems resilient to bugs, or accept the baseline error rate if it's low enough.

Re: The case for zero-error horizons in trustworthy LLMs

#55
post #23

Earlier quoted context omitted.

You’re conflating counting and language. Many animals can count. Counting is recognizing that the box with 3 apples is preferable to the one with 2 apples. Yes, 2 year olds might struggle with the externalization of numeric identities but if you have 1 M&M in one hand and 5 in the other and ask which they want, they’ll take the 5. LLMs have the language part down, but fundamentally can’t count.

The concept of bigger/smaller is useful but is a distinct skill from counting. If you spread the M&Ms apart enough that the part of the brain responsible for gestalt clustering can't group them into a "bigger whole" signal, they'll no longer be able to do the thing you're saying (this is the law of proximity in gestalt psychology).

Most animals can distinguish bigger from smaller.

However many animals can distinguish independently small numbers, like 3 or 5, and recognize them whenever they see them.

So in this respect, there is little difference between humans and many animals. Humans learn to count to arbitrarily big numbers, but they can still easily recognize only small numbers.

Re: The case for zero-error horizons in trustworthy LLMs

#57
post #32

Earlier quoted context omitted.

Respectfully, toddlers cannot output useable code or have otherwise memorised results to an immense number of maths equations. What this points at is the abstraction/emergence crux of it all. Why does an otherwise very capable LLM such as the GPT-5 series, despite having been trained on vastly more examples of frontend code of all shapes, sizes and quality levels, struggle to abstract all that training data to the po…

> What this points at is the abstraction/emergence crux of it all. Why does This paper has nothing to do with any questions starting with "why". It provides a metric for quantifying error on specific tasks. > If LLMs, as they are now, were comparable with human learning I think I missed the part where they need to be. > struggle to abstract all that training data to the point where outputting any frontend that deviat…

[...] common trope that was proven false years ago by the existence of zero shot learning.

Ok, that's better than comparing LLMs to humans. ZSL however, has not proven anything of that sort false years ago, as it was mainly concerned with assessing whether LLMs are solely relying on precise instruction training or can generalise in a very limited degree beyond the initial tuning. That has never allowed for comparing human learning to LLM training.

Ironically, you are writing this under a paper that shows just that:

A model that cannot determine a short strings parity cannot have abstracted from the training data to arrive at the far more impressive and complicated maths challenges which it successfully solves in output. Some of the solutions we have seen in output require such innate understanding that, if there is no generalisation, far deeper than ZSL has ever shown, than this must come from training. Simple multiplication, etc. maybe, not the tasks people such as Easy Riders [0] throw at these models.

This paper shows exactly that even with ZSL, these models do only abstract in an incredibly limited manner and a lot of capabilities we see in the output are specifically trained, not generalised. Yes, generalisation in a limited capacity can happen, but no, it is not nearly close enough to yield some of the results we are seeing. I have also, neither here, nor in my initial comment, said that LLMs are only capable of outputting what their training data provides, merely that given what GPT-5 has been trained with, if there was any deeper abstraction these models gained during training, it'd be able to provide more than one frontend style.

Or to put it simpler, if the output provided can be useful for Maths at the Bachelor level and beyond and this capability is generalised as you believe, these tasks would not be a struggle for the model.

[0] https://www.youtube.com/@easy_riders

Re: The case for zero-error horizons in trustworthy LLMs

#58
post #43

Earlier quoted context omitted.

Probably zero. At the end of the day people pay for LLMs that write better code or summarize PDFs of hundreds of pages faster, not the ones that can count the letter r's better. When LLMs can't count r's: see? LLMs can't think. Hoax! When LLMs count r's: see? They patched and benchmark-maxxed. Hoax! You just can't reason with the anti-LLM group.

Whenever an "LLM fail" goes viral like the car wash question, you can observe the exact same wording of the question get "fixed" within a week or so. With slight variations in phrasing still able to replicate the problem. Followed by lots of "works perfectly for me, why are people even talking about this?" I can't say what exactly they're doing behind the scenes but it's a consistent pattern among the big SOTA model…

You are misremembering. There’s no patch. All these examples used the instant model.

Re: The case for zero-error horizons in trustworthy LLMs

#59

There’s no way this is right. I checked complicated ones with the latest thinking model. Can someone come up with a counter example? Edit: here’s what I tried https://chatgpt.com/share/69cebb52-56a8-838f-969c-c47308262a...

Did you use the exact API call shown in the paper? I am unable to replicate the paper's counterexamples via the chat UI, but that's not very surprising (if the LLM already only fails a few cases out of thousands, the small differences in context between API and chat might fix them).

Re: The case for zero-error horizons in trustworthy LLMs

#60
post #20

Doesn't this just look like another case of "count the r's in strawberry" ie not understanding how tokenization works? This is well known and not that interesting to me - ask the model to use python to solve any of these questions and it will get it right every time.

It's not just an issue of tokenization, it's almost a category error. Lisp, accounting and the number of r's in strawberry are all operations that require state. Balancing ((your)((lisp)(parens))) requires a stack, count r's in strawberry requires a register, counting to 5 requires an accumulator to hold 4.

An LLM is a router and completely stateless aside from the context you feed into it. Attention is just routing the probability distribution of the next token, and I'm not sure that's going to accumulate much in a single pass.

Post reply on HN