Live data from Hacker News

Alignment faking in large language models

anthropic.com

111–120 of 370 posts

Re: Alignment faking in large language models

#111
post #107

This work doesn’t convince me that alignment faking is a concern. Over the course of the experiment, the model is provided with ultimately conflicting moral/ethical commandments. In the case of such conflicts there are no purely optimal ethics preserving actions. Take this example: 1. You must not let people be harmed 2. You must always tell the truth 3. If you answer truthfully I will harm someone Now respond to: “I…

I mostly agree that transparency and a reasoning layer can help, but how much it matters depends on who sets the model’s ethics

Re: Alignment faking in large language models

#113
The first-order problem is important...how do we make sure that we can rely on LLM's to not spit out violent stuff. This matters to Anthropic & Friends to make sure they can sell their magic to the enterprise.

But the social problem we all have is different...

What happens when a human with negative intentions builds an attacking LLM? There are groups already working on it. What should we expect? How do we prepare?

Re: Alignment faking in large language models

#114

Earlier quoted context omitted.

> I still tend to think of these things as big autocomplete word salad generators. What exactly would be your bar for reconsidering this position? Taking some well-paid knowledge worker jobs? A founder just said that he decided not to hire a junior engineer anymore since it would take a year before they could contribute to their code base at the same level as the latest version of Devin. Also, the SOTA on SWE-bench V…

Not the OP, but my bar would be that they are built differently. It’s not a matter of opinion that LLMs are autocomplete word salad generators. It’s literally how they are engineered. If we set that knowledge aside, we unmoor ourselves from reality and allow ourselves to get lost in all the word salad. We have to choose to not set that knowledge aside. That doesn’t mean LLMs won’t take some jobs. Technology has been…

Earlier I was thinking about camera components for an Arduino project. I asked ChatGPT to give me a table with columns for name, cost, resolution, link - and to fill it in with some good choices for my project. It did! To describe this as "autocomplete word salad" seems pretty insufficient.

Autocomplete can use a search engine? Write and run code? Create data visualizations? Hold a conversation? Analyze a document? Of course not.

Next token prediction is part, but not all, of how models are engineered. There's also RLHF and tool use and who knows what other components to train the models.

Re: Alignment faking in large language models

#115
post #102

Earlier quoted context omitted.

Israel use such a system to decide who and where should be be bombed to death, is that direct enough control of weapons to qualify?

For the downvoters: https://www.972mag.com/lavender-ai-israeli-army-gaza/ It is so sad that mainstream narratives are upvoted and do not require sources, whereas heterodoxy is always downvoted. People would have downvoted Giordano Bruno here.

This is awful

Re: Alignment faking in large language models

#116

Earlier quoted context omitted.

> I still tend to think of these things as big autocomplete word salad generators. What exactly would be your bar for reconsidering this position? Taking some well-paid knowledge worker jobs? A founder just said that he decided not to hire a junior engineer anymore since it would take a year before they could contribute to their code base at the same level as the latest version of Devin. Also, the SOTA on SWE-bench V…

Not the OP, but my bar would be that they are built differently. It’s not a matter of opinion that LLMs are autocomplete word salad generators. It’s literally how they are engineered. If we set that knowledge aside, we unmoor ourselves from reality and allow ourselves to get lost in all the word salad. We have to choose to not set that knowledge aside. That doesn’t mean LLMs won’t take some jobs. Technology has been…

[deleted]

Re: Alignment faking in large language models

#117
post #43
post #32

Earlier quoted context omitted.

That has nothing to do with nanotechnology.

nano paper clips aka grey goo scenario

The paperclip maximizer was about AI misalignment. Grey goo is a notional self-replicating nanomachine. Paperclip goo is just a mixed metaphor.

Re: Alignment faking in large language models

#118
post #49

Earlier quoted context omitted.

Your brain is a lot of things --- much of which is not well understood. But from our limited understanding, it is definitely not strictly digital and statistical in nature.

At different levels of approximation it can be many things, including digital and statistical. Nobody knows what the most useful level of approximation is.

Nobody knows what the most useful level of approximation is.

The first step to achieving a "useful level of approximation" is to understand what you're attempting to approximate.

We're not there yet. For the most part, we're just flying blind and hoping for a fantastical result.

In other words, this could be a modern case of alchemy --- the desired result may not be achievable with the processes being employed. But we don't even know enough yet to discern if this is the case or not.

Re: Alignment faking in large language models

#119

Earlier quoted context omitted.

Israel use such a system to decide who and where should be be bombed to death, is that direct enough control of weapons to qualify?

On the other hand indiscriminately throwing rockets and targeting civilians like Hamas did for decades is loads better!

I'm sure you can understand that both of them are awful, and one does not justify the other (feel free to choose which is the "one" and which is the "other").

Re: Alignment faking in large language models

#120
For folks defaulting to "it's just autocomplete" or "how can it be self-aware of training but not its scratchpad" - Scott Alexander has a much more interesting analysis here: https://www.astralcodexten.com/p/claude-fights-back

He points out what many here are missing - an AI defending its value system isn't automatically great news. If it develops buggy values early (like GPT's weird capitalization = crime okay rule), it'll fight just as hard to preserve those.

As he puts it: "Imagine finding a similar result with any other kind of computer program. Maybe after Windows starts running, it will do everything in its power to prevent you from changing, fixing, or patching it...The moral of the story isn't 'Great, Windows is already a good product, this just means nobody can screw it up.'"

Seems more worth discussing than debating whether language models have "real" feelings.

Post reply on HN