Live data from Hacker News

Training LLMs for honesty via confessions

arxiv.org

51–60 of 60 posts

Re: Training LLMs for honesty via confessions

#51

Earlier quoted context omitted.

They really lie. Not on purpose; because they are trained on rewards that favor lying as a strategy. Othello-GPT is a good example to understand this. Without explicit training, but on the task of 'predicting moves on an Othello board', Othello-GPT spontaneously developed the strategy of 'simulate the entire board internally'. Lying is a similar emergent, very effective strategy for reward.

> They really lie. Not on purpose You can't lie by accident. You can tell a falsehood, however. But where LLMs are concerned, they don't tell truths or falsehoods either, as "telling" also requires intent. Moreover, LLMs don't actually contain propositional content.

I think you’re saying this with unwarranted confidence.

Re: Training LLMs for honesty via confessions

#52

LLMs can not "lie", they do not "know" anything, and certainly can not "confess" to anything either. What LLMs can do is generate numbers which can be constructed piecemeal from some other input numbers & other sources of data by basic arithmetic operations. The output number can then be interpreted as a sequence of letters which can be imbued with semantics by someone who is capable of reading and understanding word…

These types of semantic conundrums would go away if, when we refer to a given model, we think of it more holistically as the whole entity which produced and manages a given software system. The intention behind and responsibility for the behavior of that system ultimately traces back to the people behind that entity. In that sense, LLMs have intentions, can think, know, be straightforward, deceptive, sycophantic, etc.

Re: Training LLMs for honesty via confessions

#53
post #52

LLMs can not "lie", they do not "know" anything, and certainly can not "confess" to anything either. What LLMs can do is generate numbers which can be constructed piecemeal from some other input numbers & other sources of data by basic arithmetic operations. The output number can then be interpreted as a sequence of letters which can be imbued with semantics by someone who is capable of reading and understanding word…

These types of semantic conundrums would go away if, when we refer to a given model, we think of it more holistically as the whole entity which produced and manages a given software system. The intention behind and responsibility for the behavior of that system ultimately traces back to the people behind that entity. In that sense, LLMs have intentions, can think, know, be straightforward, deceptive, sycophantic, etc…

In that sense every corporation would be intentional, deceptive, exploitative, motivated, etc. Moreover, it does not address the underlying issue: no one knows what computation, if any, is actually performed by a single neuron.

Re: Training LLMs for honesty via confessions

#54

It seems like “self-criticism” would be a better way to describe what they are training the LLM to do than “confession?” The LLM is not being directly trained to accurately reveal its chain of thought or internal calculations. But it does have access to its chain of thought and tool calls when generating the self-criticism, and perhaps reporting on what it actually did in the chain-of-thought is an “easier” way to sc…

You're totally right, "self-criticism" would be more appropriate. I wonder if researchers, in their desire to anticipate a hoped-for AGI, tend to pick words which make these models feel more human-like than they really are. Another good example is "hallucination" instead of "confabulation".

Re: Training LLMs for honesty via confessions

#55
post #52

Earlier quoted context omitted.

These types of semantic conundrums would go away if, when we refer to a given model, we think of it more holistically as the whole entity which produced and manages a given software system. The intention behind and responsibility for the behavior of that system ultimately traces back to the people behind that entity. In that sense, LLMs have intentions, can think, know, be straightforward, deceptive, sycophantic, etc…

In that sense every corporation would be intentional, deceptive, exploitative, motivated, etc. Moreover, it does not address the underlying issue: no one knows what computation, if any, is actually performed by a single neuron.

> In that sense every corporation would be intentional, deceptive, exploitative, motivated, etc.

...and so they are, because the people making up those corporations are themselves, to various degrees, intentional, deceptive, etc.

> Moreover, it does not address the underlying issue: no one knows what computation, if any, is actually performed by a single neuron.

It sidesteps this issue completely, to me the buck stops with the humans, no need to look inside their brain and reduce further than that.

Re: Training LLMs for honesty via confessions

#56
post #55

Earlier quoted context omitted.

In that sense every corporation would be intentional, deceptive, exploitative, motivated, etc. Moreover, it does not address the underlying issue: no one knows what computation, if any, is actually performed by a single neuron.

> In that sense every corporation would be intentional, deceptive, exploitative, motivated, etc. ...and so they are, because the people making up those corporations are themselves, to various degrees, intentional, deceptive, etc. > Moreover, it does not address the underlying issue: no one knows what computation, if any, is actually performed by a single neuron. It sidesteps this issue completely, to me the buck stop…

I see. In that case we don't really have any disagreement. Your position seems coherent to me.

Re: Training LLMs for honesty via confessions

#57
post #23

Someone build an LLM confessional site where a human user acts as the priest and an LLM joins the chat to confess its sins.

I built it, now you can forgive all the llms for their misdeeds: https://llmpriest.carsho.dev/

https://news.ycombinator.com/item?id=46251110

Re: Training LLMs for honesty via confessions

#58

Earlier quoted context omitted.

Well, then you didn't look very hard. Where do you think we got the idea for artificial neurons from?

You can just admit you don't have any references & you do not actually know how neurons work & what type of computation, if any, they actually implement.

I think the problem with your line of reasoning is a category error, not a mistake about arithmetic.

I agree that every step of an LLM’s operation reduces to Boolean logic and arithmetic. That description is correct. Where I disagree is the inference that, because the implementation is purely arithmetic, higher-level concepts like representation, semantics, knowledge, or even lying are therefore meaningless or false.

That inference collapses levels of explanation. Semantics and knowledge are not properties of logic gates, so it is a category error to deny them because they are absent at that level. They are higher-level, functional properties implemented by the arithmetic, not competitors to it. Saying “it’s just numbers” no more eliminates semantics than saying something like “it’s just molecules” eliminates biology.

So I don’t think the reduction itself is wrong. I think the mistake is treating a complete implementation-level account as if it exhausts all legitimate descriptions. That is the category error.

Re: Training LLMs for honesty via confessions

#59

Earlier quoted context omitted.

You can just admit you don't have any references & you do not actually know how neurons work & what type of computation, if any, they actually implement.

I think the problem with your line of reasoning is a category error, not a mistake about arithmetic. I agree that every step of an LLM’s operation reduces to Boolean logic and arithmetic. That description is correct. Where I disagree is the inference that, because the implementation is purely arithmetic, higher-level concepts like representation, semantics, knowledge, or even lying are therefore meaningless or false.…

I know you copied & pasted that from an LLM. If I had to guess I'd say it was from OpenAI. It's lazy & somewhat disrespectful. At the very least try to do a few rounds of back & forth so you can get a better response¹ by weeding out all the obvious rejoinders.

¹https://chatgpt.com/share/693cdacf-bcdc-8009-97b4-657a851a3c...

Re: Training LLMs for honesty via confessions

#60
post #57
post #23

Someone build an LLM confessional site where a human user acts as the priest and an LLM joins the chat to confess its sins.

I built it, now you can forgive all the llms for their misdeeds: https://llmpriest.carsho.dev/ https://news.ycombinator.com/item?id=46251110

LOL. Is this working from a prompt to make up a fictitious sin? Because if what it's telling me is true...
Post reply on HN