Live data from Hacker News

Alignment faking in large language models

anthropic.com

181–190 of 370 posts

Re: Alignment faking in large language models

#181
post #102

Earlier quoted context omitted.

Israel use such a system to decide who and where should be be bombed to death, is that direct enough control of weapons to qualify?

For the downvoters: https://www.972mag.com/lavender-ai-israeli-army-gaza/ It is so sad that mainstream narratives are upvoted and do not require sources, whereas heterodoxy is always downvoted. People would have downvoted Giordano Bruno here.

It's mainstream enough to have a Wikipedia article.

https://en.wikipedia.org/wiki/AI-assisted_targeting_in_the_G...

Re: Alignment faking in large language models

#182

Earlier quoted context omitted.

You're comparing the actions of what most people here view as a democratic state (parlamentary republic) and an opaquely run terrorist organization. We're talking about potential consequences of giving AIs influence on military decisions. To that point, I'm not sure what your comment is saying. Is it perhaps: "we're still just indiscriminately killing civilians just as always, so giving AI control is fine"?

> we're still just indiscriminately killing civilians just as always, so giving AI control is fine I don't even want to respond because the "we" and "as always" here is doing a lot . I don't have it in me to have an extended discussion to address how indiscriminately killing civilians was never accepted practice in modern warfare. Anyways. There are two conditions in which I see this argument(?) is useful. If you ass…

> I don't have it in me to have an extended discussion to address how indiscriminately killing civilians was never accepted practice in modern warfare.

I did not claim it is/was accepted practice. I was asking if "doing it with AI is just the same so what's the big deal" was your position on the general issue (of AI making decisions in war), which I thought was a possible interpretation of your previous comment.

> No, I replied to a comment that was talking about a specific example.

OK. That means the two of us were/are just talking past each other and won't be having an interesting discussion.

Re: Alignment faking in large language models

#183

Earlier quoted context omitted.

> Restricting access to information is not really doing much at all. Why not? Restricting access to information is of course harder but that's no argument for it not doing anything. Governments restrict access to "state secrets" all the time. Depending on the topic, it's hard but may still be effective and worth it. For example, you seem to agree that restricting access to weapons makes sense. What to do about 3D-pri…

Meh, 3D printed guns are a stupid example that gets trotted out just because it sounds futuristic. In WW2 you had many examples of machinists in occupied Europe who produced workable submachine guns - far better than any 3D-printed firearm - right under the nose of the Nazis. Literally when armed soldiers could enter your house and inspect it at any time. Our machining tools today are much better, but no-one is conce…

People actually are concerned about homemade pistols and SMGs being used by criminals, though. It comes up quite often in Europe these days, especially in UK.

And, yes, in principle, 3D printing doesn't really bring anything new to the table since you could always machine a gun, and the tools to do so are all available. The difference is ease of use - 3D printing lowered the bar for "people competent enough to manufacture dangerous things" enough that your latter argument no longer applies.

FWIW I don't know the answer to OP's question even so. I don't think we should be banning 3D printed gun designs, or, for that matter, that even if we did, such a ban would be meaningfully enforceable. I don't think 3D printers should be banned, either. This feels like one of those cases where you have to accept that new technology has some unfortunate side effects.

Re: Alignment faking in large language models

#184
post #130

Earlier quoted context omitted.

Indeed. If the smart lawnmower (Powered by AI™, as seen on television) decides that not being turned off is the best way to achieve its ultimate goal of getting your lawn mowed, it doesn't matter whether the completely unnecessary LLM inside is just a dumb copyright infrigement machine and probably just copying the plot it learned in some sci-fi story somewhere in training set. Your foot is still getting mowed! AIs d…

IF one maintains a clear understanding of how the technology actually works, THEN one will make good decisions about whether to put it charge of the lawnmower in the first place. Anthropic is in the business of selling AI. Of course they are going to approach alignment as a necessary and solvable problem. The rest of us don’t have to go along with that, though. Why is it even necessary to use an LLM to mow a lawn? Th…

> Why is it even necessary to use an LLM to mow a lawn? There is more to AI than generative LLMs.

As if this reasoning will stop people?

On the front page today we have:

\> Our system combines a robotic platform (we call the first one Maurice) with an AI agent that understands the environment, plans actions, and executes them using skills you've taught it or programmed within our SDK.

https://news.ycombinator.com/item?id=42451707

Re: Alignment faking in large language models

#185
post #130

Earlier quoted context omitted.

Indeed. If the smart lawnmower (Powered by AI™, as seen on television) decides that not being turned off is the best way to achieve its ultimate goal of getting your lawn mowed, it doesn't matter whether the completely unnecessary LLM inside is just a dumb copyright infrigement machine and probably just copying the plot it learned in some sci-fi story somewhere in training set. Your foot is still getting mowed! AIs d…

IF one maintains a clear understanding of how the technology actually works, THEN one will make good decisions about whether to put it charge of the lawnmower in the first place. Anthropic is in the business of selling AI. Of course they are going to approach alignment as a necessary and solvable problem. The rest of us don’t have to go along with that, though. Why is it even necessary to use an LLM to mow a lawn? Th…

LOL

smart fridge anyone?

Re: Alignment faking in large language models

#186
post #132

Earlier quoted context omitted.

Exactly. The discussion is going to change real fast when LLMs are wrapped in some sort OODA loop type thing and crammed into some sort of humanoid robot that carries hedge trimmers.

why would you want to let a LLM have any agentic interface to the real world though

Fantastic question, probably best-answered by these folks!

https://docs.innate.bot/docs.innate.bot

Re: Alignment faking in large language models

#187
post #146

Earlier quoted context omitted.

I'd be very interested in an example of someone who said "it's just autocomplete" and also explicitly brought up the risk of something like alignment faking, before 2024. I can think of examples of people who've been talking about this kind of thing for years, but they're all people who have no trouble with applying the adjective "intelligent" to models.

If you go back through my Hacker News comments, I believe you'll see this. Perhaps look for keywords "GPT-2", "prediction", and "agent". (I don't know how to search HN comments efficiently.) I was talking about this sort of thing in 2018, though I don't think I published anything that's still accessible, and I'd hardly call myself an expert: it's just obviously how the system works.

Searching HN comments is probably easiest done through Algolia:

https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...

Re: Alignment faking in large language models

#188

Earlier quoted context omitted.

Your brain isn't a truth machine. It can't be, it has to create an inner map that relates to the outer world. You have never seen the real world. You are calculating the angular distance between signals just like Claude is. It's more of a question of degree than category.

Claude hasn't been trained with skin in the game. That is one of the reasons it confabulates so readily. The weights and biases are shaped by an external classification. There isn't really a way to train consequences into the model like natural selection has been able to train us.

I would say that consequences are exactly the modification of weights and biases when models make a mistake.

How many of us take a Machiavellian approach to calculate the combined chance of getting caught and the punishment if we are, instead of just going with gut feelings based on our internalised model from a lifetime of experience? Some, but not most of us.

What we get from natural selection is instinct, which I think includes what smiles look like, but that's just a fast way to get feedback.

Re: Alignment faking in large language models

#189

Earlier quoted context omitted.

Meh, 3D printed guns are a stupid example that gets trotted out just because it sounds futuristic. In WW2 you had many examples of machinists in occupied Europe who produced workable submachine guns - far better than any 3D-printed firearm - right under the nose of the Nazis. Literally when armed soldiers could enter your house and inspect it at any time. Our machining tools today are much better, but no-one is conce…

There's very little competence (and also money) required to buy a 3D printer, download a design and print it. A lot less competence than "being a machinist". The point is that making dangerous things is becoming a lot easier over time.

you can 3-D print parts of a gun but the important parts are still metal which you need to machine. I’m not sure how much easier you just made it … if someone’s making a gun in their basement are you really concerned whether it takes 20 hours or 10? What you should be really concerned about is when the cost of milling machines comes down, which is happening, quick, make them illegal

Re: Alignment faking in large language models

#190
post #154
post #146

Earlier quoted context omitted.

I'd be very interested in an example of someone who said "it's just autocomplete" and also explicitly brought up the risk of something like alignment faking, before 2024. I can think of examples of people who've been talking about this kind of thing for years, but they're all people who have no trouble with applying the adjective "intelligent" to models.

The point is that if the limitations of current LLMs persist, regardless of how much better they get, this is not a problem at all, or at least not a new one . Let's say you are given the declaration but not the implementation of a function with the following prototype: const char * AskTheLLM(const char *prompt); Putting this function in charge of anything, unless a restricted interface is provided so that it can't d…

> Let's say you are given the declaration but not the implementation of a function with the following prototype:

> const char * AskTheLLM(const char prompt);

> Putting this function in charge of anything, unless a restricted interface is provided so that it can't do much damage, is simply terrible engineering and not at all how anything is done.

Yes, but that's exactly how people* use any system that has an air of authority, unless they're being very careful to apply critical thinking and skepticism. It's why confidence scams and advertising work.

This is also at the heart of current "alignment" practices. The goal isn't so much to have a model that can't automate harm as it is to have one that won't provide authoritative-sounding but "bad" answers to people who might believe them. "Bad," of course, covers everything from dangerously incorrect to reputational embarrassments.

Post reply on HN