Earlier quoted context omitted.
They really lie. Not on purpose; because they are trained on rewards that favor lying as a strategy. Othello-GPT is a good example to understand this. Without explicit training, but on the task of 'predicting moves on an Othello board', Othello-GPT spontaneously developed the strategy of 'simulate the entire board internally'. Lying is a similar emergent, very effective strategy for reward.
> They really lie. Not on purpose You can't lie by accident. You can tell a falsehood, however. But where LLMs are concerned, they don't tell truths or falsehoods either, as "telling" also requires intent. Moreover, LLMs don't actually contain propositional content.
Training LLMs for honesty via confessions
51–60 of 60 posts
Re: Training LLMs for honesty via confessions
#52LLMs can not "lie", they do not "know" anything, and certainly can not "confess" to anything either. What LLMs can do is generate numbers which can be constructed piecemeal from some other input numbers & other sources of data by basic arithmetic operations. The output number can then be interpreted as a sequence of letters which can be imbued with semantics by someone who is capable of reading and understanding word…
Re: Training LLMs for honesty via confessions
#53LLMs can not "lie", they do not "know" anything, and certainly can not "confess" to anything either. What LLMs can do is generate numbers which can be constructed piecemeal from some other input numbers & other sources of data by basic arithmetic operations. The output number can then be interpreted as a sequence of letters which can be imbued with semantics by someone who is capable of reading and understanding word…
These types of semantic conundrums would go away if, when we refer to a given model, we think of it more holistically as the whole entity which produced and manages a given software system. The intention behind and responsibility for the behavior of that system ultimately traces back to the people behind that entity. In that sense, LLMs have intentions, can think, know, be straightforward, deceptive, sycophantic, etc…
Re: Training LLMs for honesty via confessions
#54It seems like “self-criticism” would be a better way to describe what they are training the LLM to do than “confession?” The LLM is not being directly trained to accurately reveal its chain of thought or internal calculations. But it does have access to its chain of thought and tool calls when generating the self-criticism, and perhaps reporting on what it actually did in the chain-of-thought is an “easier” way to sc…
Re: Training LLMs for honesty via confessions
#55Earlier quoted context omitted.
These types of semantic conundrums would go away if, when we refer to a given model, we think of it more holistically as the whole entity which produced and manages a given software system. The intention behind and responsibility for the behavior of that system ultimately traces back to the people behind that entity. In that sense, LLMs have intentions, can think, know, be straightforward, deceptive, sycophantic, etc…
In that sense every corporation would be intentional, deceptive, exploitative, motivated, etc. Moreover, it does not address the underlying issue: no one knows what computation, if any, is actually performed by a single neuron.
...and so they are, because the people making up those corporations are themselves, to various degrees, intentional, deceptive, etc.
> Moreover, it does not address the underlying issue: no one knows what computation, if any, is actually performed by a single neuron.
It sidesteps this issue completely, to me the buck stops with the humans, no need to look inside their brain and reduce further than that.
Re: Training LLMs for honesty via confessions
#56Earlier quoted context omitted.
In that sense every corporation would be intentional, deceptive, exploitative, motivated, etc. Moreover, it does not address the underlying issue: no one knows what computation, if any, is actually performed by a single neuron.
> In that sense every corporation would be intentional, deceptive, exploitative, motivated, etc. ...and so they are, because the people making up those corporations are themselves, to various degrees, intentional, deceptive, etc. > Moreover, it does not address the underlying issue: no one knows what computation, if any, is actually performed by a single neuron. It sidesteps this issue completely, to me the buck stop…
Re: Training LLMs for honesty via confessions
#57Someone build an LLM confessional site where a human user acts as the priest and an LLM joins the chat to confess its sins.
Re: Training LLMs for honesty via confessions
#58Earlier quoted context omitted.
Well, then you didn't look very hard. Where do you think we got the idea for artificial neurons from?
You can just admit you don't have any references & you do not actually know how neurons work & what type of computation, if any, they actually implement.
I agree that every step of an LLM’s operation reduces to Boolean logic and arithmetic. That description is correct. Where I disagree is the inference that, because the implementation is purely arithmetic, higher-level concepts like representation, semantics, knowledge, or even lying are therefore meaningless or false.
That inference collapses levels of explanation. Semantics and knowledge are not properties of logic gates, so it is a category error to deny them because they are absent at that level. They are higher-level, functional properties implemented by the arithmetic, not competitors to it. Saying “it’s just numbers” no more eliminates semantics than saying something like “it’s just molecules” eliminates biology.
So I don’t think the reduction itself is wrong. I think the mistake is treating a complete implementation-level account as if it exhausts all legitimate descriptions. That is the category error.
Re: Training LLMs for honesty via confessions
#59Earlier quoted context omitted.
You can just admit you don't have any references & you do not actually know how neurons work & what type of computation, if any, they actually implement.
I think the problem with your line of reasoning is a category error, not a mistake about arithmetic. I agree that every step of an LLM’s operation reduces to Boolean logic and arithmetic. That description is correct. Where I disagree is the inference that, because the implementation is purely arithmetic, higher-level concepts like representation, semantics, knowledge, or even lying are therefore meaningless or false.…
¹https://chatgpt.com/share/693cdacf-bcdc-8009-97b4-657a851a3c...
Re: Training LLMs for honesty via confessions
#60Someone build an LLM confessional site where a human user acts as the priest and an LLM joins the chat to confess its sins.
I built it, now you can forgive all the llms for their misdeeds: https://llmpriest.carsho.dev/ https://news.ycombinator.com/item?id=46251110