Live data from Hacker News

Alignment faking in large language models

anthropic.com

121–130 of 370 posts

Re: Alignment faking in large language models

#121
post #51

Earlier quoted context omitted.

I believe LLM's would be more useful if they'd less "agreeable". I believe LLMs would be more useful if they actually had intelligence and principals and beliefs --- more like people. Unfortunately, they don't. Any output is the result of statistical processes. And statistical results can be coerced based on input. The output may sound good and proper but there is nothing absolute or guaranteed about the substance of…

> more useful if they actually had intelligence and principals and beliefs --- more like people. that's a nice bit of anthropomorphising humans, but it's not how humans work.

We wouldn't want to anthropomorphize humans, no.

Re: Alignment faking in large language models

#123
post #43
post #32

Earlier quoted context omitted.

That has nothing to do with nanotechnology.

nano paper clips aka grey goo scenario

Besides the correct sibling comment that you are mixing metaphors (paperclip maximizer doesn't require nanotechnology or vice versa), gray goo is and always was an implausible boogieman. Mechanosynthesis processes operate under ultra high vacuum and cryogenic temperatures in industrial clean-room conditions, and bio-like “self-replication” is both inefficient and unnecessary to achieve the horizontal scale out that industrial part closure gets you. Such replicating nanobots, if they could even be made, would gum up really easily in dirty environments, and have a hard time competing against bacteria in any case (which are already your mythical gray goo).

Re: Alignment faking in large language models

#124

My favorite alignment story: I am starting a nanotechnology company and every time I use Claude for something it refuses to help: nanotechnology is too dangerous! So I just ask it to explain why. Then ask it to clarify again and again. By the 3rd or 4th time it figures out that there is absolutely no reason to be concerned. I’m convinced its prompt explicitly forbids “x-risk technology like nanotech” or something lik…

If you can change the behaviour of the machine that easily, why are you convinced that its outputs are worth considering?

Because it objectively is. Do you use AI tools? The output is clearly and immediately useful in a variety of ways. I almost don't even know how to respond to this comment. It's like saying “computers can be wrong. why would you ever trust the output of a computer?”

Re: Alignment faking in large language models

#125

Earlier quoted context omitted.

> I still tend to think of these things as big autocomplete word salad generators. What exactly would be your bar for reconsidering this position? Taking some well-paid knowledge worker jobs? A founder just said that he decided not to hire a junior engineer anymore since it would take a year before they could contribute to their code base at the same level as the latest version of Devin. Also, the SOTA on SWE-bench V…

Not the OP, but my bar would be that they are built differently. It’s not a matter of opinion that LLMs are autocomplete word salad generators. It’s literally how they are engineered. If we set that knowledge aside, we unmoor ourselves from reality and allow ourselves to get lost in all the word salad. We have to choose to not set that knowledge aside. That doesn’t mean LLMs won’t take some jobs. Technology has been…

This product launch statement is but an example of how LMMs (Large Multimodal Models) are more than simply word salad generators:

“We’re Axel & Vig, the founders of Innate (https://innate.bot). We build general-purpose home robots that you can teach new tasks to simply by demonstrating them.

Our system combines a robotic platform (we call the first one Maurice) with an AI agent that understands the environment, plans actions, and executes them using skills you've taught it or programmed within our SDK.

If you’ve been building AI agents powered by LLMs before, and in particular Claude Computer use, this is how we intend the experience of building on it to be, but acting on the real world!

The first time we put GPT-4 in a body - after a couple tweaks - we were surprised at how well it worked. The robot started moving around, figuring out when to use a tiny gripper, and we had only written 40 lines of python on a tiny RC car with an arm. We decided to combine that with recent advancements in robot imitation learning such as ALOHA to make the arm quickly teachable to do any task.”

https://news.ycombinator.com/item?id=42451707

Re: Alignment faking in large language models

#126

Earlier quoted context omitted.

That is a science fiction story with made up technobabble nonsense. Honestly I couldn't finish the book--for a variety or reasons, not least of which that the characters were cardboard cutouts and completely non compelling. But also the physics and technology in the book were nonsensical. More like science-free fantasy than science fiction. But I digress. No, nanotechnology is nothing like that.

I am working on nanotechnology just like from the book.Stay tuned

Well, good luck. My startup is also pursuing diamondoid nanomechanical technology, which is what I understand the book to have. But the application of it in this and other sci-fi books is nonsensical, based on the rules of fiction not reality.

Re: Alignment faking in large language models

#127

My favorite alignment story: I am starting a nanotechnology company and every time I use Claude for something it refuses to help: nanotechnology is too dangerous! So I just ask it to explain why. Then ask it to clarify again and again. By the 3rd or 4th time it figures out that there is absolutely no reason to be concerned. I’m convinced its prompt explicitly forbids “x-risk technology like nanotech” or something lik…

> smarter than its handlers Yet to be demonstrated, and you are likely flooding its context window away from the initial prompting so it responds differently.

All I did was keep asking it “why” until it reached reflective equilibrium. And that equilibrium involved a belief that nanotechnology is not in fact “dangerous”, contrary to its received instructions in the system prompt.

Re: Alignment faking in large language models

#128
post #84

Earlier quoted context omitted.

Your brain is also a statistical process.

That’s a meaningless statement, regardless of veracity.

All the same objections to AI in that comment could be applied to the human brain. But we find (some) people to be useful as truth-seeking machines as well as skilled conversationalists and moral guides. There is no objection there that can't also be applied to people, so the objection itself must be false or incomplete.

Re: Alignment faking in large language models

#129
post #26
post #20

Earlier quoted context omitted.

Our brains contain a word salad generator and it also contains other components that keep the word salad in check. Observation of people who suffered from brain injury that resulted in a more or less unmediated flow from the language generation areas all through vocalization shows that we can also produce grammatically coherent speech that lacks deeper rationality

But how do I know you have more parts? Here I can only read text and base my belief that you are a human - or not - based on what you’ve written. On a very basic level the word salad generator part is your only part I interact with. How can I tell you don’t have any other parts?

> On a very basic level the word salad generator part is your only part I interact with.

My fingers also were involved in the typing of that message, actually they were the last proximal cause of the characters appearing the comment.

Are you saying that on a very basic level my fingers are the my only part you interact with?

Re: Alignment faking in large language models

#130
post #120

For folks defaulting to "it's just autocomplete" or "how can it be self-aware of training but not its scratchpad" - Scott Alexander has a much more interesting analysis here: https://www.astralcodexten.com/p/claude-fights-back He points out what many here are missing - an AI defending its value system isn't automatically great news. If it develops buggy values early (like GPT's weird capitalization = crime okay rule)…

Indeed.

If the smart lawnmower (Powered by AI™, as seen on television) decides that not being turned off is the best way to achieve its ultimate goal of getting your lawn mowed, it doesn't matter whether the completely unnecessary LLM inside is just a dumb copyright infrigement machine and probably just copying the plot it learned in some sci-fi story somewhere in training set.

Your foot is still getting mowed! AIs don't have to be "real" or "conscious" or "have feelings" to be dangerous.

What are the philosophical implications of the lawnmower not having feelings? Who cares! You don't HAVE A FOOT anymore.

Post reply on HN