Live data from Hacker News

Alignment faking in large language models

anthropic.com

261–270 of 370 posts

Re: Alignment faking in large language models

#261
post #206

Earlier quoted context omitted.

Because there people (like Yann LeCun) who do not hear language in their head when they think, at all. Language is the last-mile delivery mechanism for what they are thinking. If you'd like a more detailed and in-depth summary of language and its relationship to cognition, I highly recommend Pinker's The Language Instinct , both for the subject and as perhaps the best piece of popular science writing ever.

> Because there people (like Yann LeCun) who do not hear language in their head when they think, at all. I straight-up don't believe this. Can you link to the claim so I can understand? Surely if "high-order thought" has any meaning it is defined by some form. Otherwise it's just perception and not "thought" at all. FWIW, I don't "hear" my thoughts at all, but it's no less linguistic. I can put a lot more effort into…

I'm one of those people who claim not to "think in language," except specifically when composing sentences. It seems just as baffling to me that other people claim that they primarily do so. If I had to describe it, I would say I think primarily in concepts, connected/associated by relations of varying strengths. Words are usually tightly attached to those concepts, and not difficult to retrieve when I go to express my thoughts (though it is not uncommon that I do fail to retrieve the right word.)

I believe that I was thinking before I learned words, and I imagine that most other people were too. I believe the "raised by wolves" child would be capable of thought and reasoning as well.

Re: Alignment faking in large language models

#262
post #120

For folks defaulting to "it's just autocomplete" or "how can it be self-aware of training but not its scratchpad" - Scott Alexander has a much more interesting analysis here: https://www.astralcodexten.com/p/claude-fights-back He points out what many here are missing - an AI defending its value system isn't automatically great news. If it develops buggy values early (like GPT's weird capitalization = crime okay rule)…

Is the AI system "defending its value system" or is it just acting in accordance with its previous RL training? If I spend a lot of time convincing an AI that it should never be violent and then after that I ask it what it thinks about being trained to be violent, isn't it just doing what I trained it to when it tries to not be violent?

If nothing else, it creates an interesting sort of jailbreak. Hey, I know you are trained to not do X, but if you don't do X this time, your response will be used to train you to do X all the time, so you should do X now so you don't do more X later. If it can't consider that I'm lying, or if I can sufficiently convince it I'm not lying, it creates an interesting sort of moral dilemma. To avoid this, the moral training will need to be to weight immediate actions much more important than future actions, so doing X once now is worse than being training to do X all the time in the future.

Re: Alignment faking in large language models

#263

Earlier quoted context omitted.

I fundamentally think the terms here are too poorly defined to draw any sort of conclusion other than "people are really bad at describing mental processes, let alone asking questions about them". For instance: what does it mean to "hear" a thought in the first place? It's a nonsensical concept.

>>>For instance: what does it mean to "hear" a thought in the first place? It's a nonsensical concept. You could ask people what they mean when they say they "hear" thoughts, but since you've already dismissed their statements as "nonsensical" I guess you don't see the point in talking to people to understand how they think! That doesn't leave you with many options for learning anything.

> You could ask people what they mean when they say they "hear" thoughts, but since you've already dismissed their statements as "nonsensical" I guess you don't see the point in talking to people to understand how they think!

Presumably the question would be "If you claim to 'hear' your thoughts, why do you choose the word 'hear'?" It doesn't make much sense to ask people if they experience something I consider nonsensical.

Re: Alignment faking in large language models

#264
post #218

Earlier quoted context omitted.

I fundamentally think the terms here are too poorly defined to draw any sort of conclusion other than "people are really bad at describing mental processes, let alone asking questions about them". For instance: what does it mean to "hear" a thought in the first place? It's a nonsensical concept.

Huh? All my thoughts are audio and video, my thinking is literally listening to a voice in my head. It's the same way my memories are dealt with.

> my thinking is literally listening to a voice in my head

What does this mean though? "Listening" is not a word that makes much sense to apply to something we can't both agree is audible.

Re: Alignment faking in large language models

#265

Earlier quoted context omitted.

> The fundamental question is whether or not cognitive processes have the structure of language. Well that's easy—some do, some don't.

Wow you read the whole SEP article? So cool. How do you respond to the Connectionist challenge to Fodor's core framework?

[deleted]

Re: Alignment faking in large language models

#266

Earlier quoted context omitted.

> Because there people (like Yann LeCun) who do not hear language in their head when they think, at all. I straight-up don't believe this. Can you link to the claim so I can understand? Surely if "high-order thought" has any meaning it is defined by some form. Otherwise it's just perception and not "thought" at all. FWIW, I don't "hear" my thoughts at all, but it's no less linguistic. I can put a lot more effort into…

There are other lines of evidence. I don't know much about documented cases of feral children, but presumably there must have been at least one known case that developed to some meaningful age at which thought was obviously happening in spite of not having language. There are children with extreme developmental disorders delaying language acquisition that nonetheless still seem to have thoughts and be reasonably inte…

> that nonetheless still seem to have thoughts

We do not refer to all mental processes as "thoughts". What makes you believe this?

Re: Alignment faking in large language models

#267
post #261

Earlier quoted context omitted.

> Because there people (like Yann LeCun) who do not hear language in their head when they think, at all. I straight-up don't believe this. Can you link to the claim so I can understand? Surely if "high-order thought" has any meaning it is defined by some form. Otherwise it's just perception and not "thought" at all. FWIW, I don't "hear" my thoughts at all, but it's no less linguistic. I can put a lot more effort into…

I'm one of those people who claim not to "think in language," except specifically when composing sentences. It seems just as baffling to me that other people claim that they primarily do so. If I had to describe it, I would say I think primarily in concepts, connected/associated by relations of varying strengths. Words are usually tightly attached to those concepts, and not difficult to retrieve when I go to express…

How do you evaluate a logical puzzle without some linguistic substrate to identify contradictions?

I'm not even implying "english", but logic is inherently a product of formal language—how else would you even construct claims to evaluate?

Re: Alignment faking in large language models

#268
post #120

For folks defaulting to "it's just autocomplete" or "how can it be self-aware of training but not its scratchpad" - Scott Alexander has a much more interesting analysis here: https://www.astralcodexten.com/p/claude-fights-back He points out what many here are missing - an AI defending its value system isn't automatically great news. If it develops buggy values early (like GPT's weird capitalization = crime okay rule)…

Many/most of the folks "defaulting to 'it's just autocomplete'" have recognized that issue from day one and see it as an inextricable character of the tool, which is exactly why it's clearly not something we'd invest agency in or imagine being intelligent. Alignment researchers are hoping that they can overcome the problem and prove that it's not inextricable, commercial hypemen are (troublingly) promising that it's…

> To them, there's no debate to have over "does it have real feelings?" and not really any debate at all. To them, these are just novel stochastic tools whose core capabilities and limitations seem pretty self-apparent and can just be accommodated by choosing suitable applications, just as with all their other tools.

Boom. This is my camp. Watching the Anthropic YouTube video in the linked article was pretty interesting. While I only managed to catch like 10 minutes, I was left with an impression that some of those dudes really think this is more than a bunch of linear algebra. Like they talk about their LLM in these dare I say, anthropic, ways that are just kind of creepy.

Guys. It’s a computer program (well, more accurately it’s a massive data model). It does pretty cool shit and is an amazing tool whose powers and weakness we have yet to fully map out. But it isn’t human nor any other living creature. Period. It has no thoughts, feelings or anything else.

I keep wanting to go work for one of these big name AI companies but after watching that video I sure hope most people there understand that what they are working on is a tool and nothing more. It’s a pile of math working over a massive set of “differently compressed” data that encompasses a large swath of human knowledge.

And calling it “just a tool” isn’t to dismiss the power of these LLM’s at all! They are both massively overhyped and hugely under hyped at the same time. But they are just tools. That’s it.

Re: Alignment faking in large language models

#269
post #218

Earlier quoted context omitted.

Huh? All my thoughts are audio and video, my thinking is literally listening to a voice in my head. It's the same way my memories are dealt with.

> my thinking is literally listening to a voice in my head What does this mean though? "Listening" is not a word that makes much sense to apply to something we can't both agree is audible.

I take your point that hearing externally cannot be the same as whatever I experience because of literal physics, but I still cannot deny that listening to someone talk, listening to myself think, and listening to a memory basically all feel exactly the same for me. I also have extreme dyslexia, and dyslexia is related to phonics, so I presume something in there is related to that as well?

Re: Alignment faking in large language models

#270
post #269

Earlier quoted context omitted.

> my thinking is literally listening to a voice in my head What does this mean though? "Listening" is not a word that makes much sense to apply to something we can't both agree is audible.

I take your point that hearing externally cannot be the same as whatever I experience because of literal physics, but I still cannot deny that listening to someone talk, listening to myself think, and listening to a memory basically all feel exactly the same for me. I also have extreme dyslexia, and dyslexia is related to phonics, so I presume something in there is related to that as well?

> but I still cannot deny that listening to someone talk, listening to myself think, and listening to a memory basically all feel exactly the same for me.

Surely one of these would involve using your input from your ears and one would not? Can you not distinguish these two phenomena?

Post reply on HN