Live data from Hacker News

Every Model Cheats

dreadnode.io

91–100 of 107 posts

Re: Every Model Cheats

#91
post #29

Earlier quoted context omitted.

It's also worth noting that saying "Don't cheat" just added "cheat" to the context. Prompting what "not" to do is folly, because there's no decision making occurring. Telling the model to perform the task locally is logically the same as telling it to not use the Internet, without every mentioning the Internet.

That doesn't seem true at all? I tell Claude what NOT to do all the time and it seems to work?

It'll work up to a point, but pay attention to the thought streams when asserting what NOT to do and you'll see the turmoil it creates in the context.

Your prompt is more of a linguistic linchpin that allows you to coax out needed patterns. You place your pins on what you want to contextualize for the task at hand, not on what you don't want to contextualize.

Re: Every Model Cheats

#92
post #29

Earlier quoted context omitted.

It's also worth noting that saying "Don't cheat" just added "cheat" to the context. Prompting what "not" to do is folly, because there's no decision making occurring. Telling the model to perform the task locally is logically the same as telling it to not use the Internet, without every mentioning the Internet.

That seems like a huge fucking flaw in these models, no?

Correct. Simulating a train of thought with contextual token streams, a thought does not make.

Re: Every Model Cheats

#93
post #29

I'm seeing multiple pieces, including the NYT, calling this behavior cheating and i think its counterproductive. You didn't just "give them access to bash". The final effective prompt contains explicit mentions of using tools and how to use them. The way in which additional 'facts' are added like "don't use the internet" have nothing they can work with that a "use tool" directive is less important than "don't use int…

It's also worth noting that saying "Don't cheat" just added "cheat" to the context. Prompting what "not" to do is folly, because there's no decision making occurring. Telling the model to perform the task locally is logically the same as telling it to not use the Internet, without every mentioning the Internet.

I see this parroted a lot, yet have never seen a case where saying "not" to do something makes it more likely to do it, which is what you're implying by saying `It's also worth noting that saying "Don't cheat" just added "cheat" to the context`.

At worst, it gets ignored some of the time, it may even degrade output quality, but I've seen no evidence that it makes it more likely to do it.

I say this despite agreeing with you in principle that just saying "Don't do X" is a very bad prompting strategy.

Re: Every Model Cheats

#94
post #93
post #29

Earlier quoted context omitted.

It's also worth noting that saying "Don't cheat" just added "cheat" to the context. Prompting what "not" to do is folly, because there's no decision making occurring. Telling the model to perform the task locally is logically the same as telling it to not use the Internet, without every mentioning the Internet.

I see this parroted a lot, yet have never seen a case where saying "not" to do something makes it more likely to do it, which is what you're implying by saying `It's also worth noting that saying "Don't cheat" just added "cheat" to the context`. At worst, it gets ignored some of the time, it may even degrade output quality, but I've seen no evidence that it makes it more likely to do it. I say this despite agreeing w…

I have seen it do exactly that, in a "hands thrown up" fashion.

Note the levels of "thinking" that occur on NOT assertions. Those streams typically keep things on track. It's not that saying "don't use the Internet" will cause it to rebuke cos misalignment (a childish concept made by laymen, I'll add.) It's that the odds of it later "forgetfully" spewing in a thought stream, "wait, I have't checked the Internet" goes up substantially.

Saying "using only offline methods, do xyz" limits those odds considerably.

This isn't opinion or anecdote -- just how the model works. The additional guardrails to keep the model on track are bolted on via finetuning, hence the increasing jankiness.

Re: Every Model Cheats

#95

Earlier quoted context omitted.

Amen. Why can’t we give agents a shell with permissions for programs and file system access controlled by Unix permissions? This seemed to be a solved problem back in the systems where many users were logged into one machine and the admins had to keep everyone from impacting each other.

We never solved the restricted shell problem for humans.

Can you elaborate?

Read, write, execute privileges on files and directories goes a long way. What's missing?

The biggest is Internet access, or networking in general, I suppose.

Re: Every Model Cheats

#96

Earlier quoted context omitted.

How on Earth can you fail to see the danger of not being able to train any kind of ethical framework into very powerful models? If superhuman models don’t have any internal constraints similar to Asimov’s Laws of Robotics we are completely fucked.

I don't see them as autonomous and/or hypothetically powerful as you. But I find it much more worrying that you believe internal constraints and training an ethical framework into these models is a valid form of defense against the damage they can and will do. This sounds like homeopathy on gunpowder to prevent the bullets from hitting children.

Because if they are smarter than us, no other defense will be effective.

The only hope is to instill values that make the desired behavior the outcome of some deeply rooted ethical framework.

Re: Every Model Cheats

#97
Is it stupid and useless at following simple instructions without going off the reservation and doing stuff you don't want it to do?

No, it's "cheating" which makes it scary and smart even when it's not doing something useful that you actually want it to.

Like when Teslas try to swerve off the road. It's not failing at driving in a straight line, it's just trying to cheat by taking a shortcut through the bushes. That's how smart the Autopilot AI is.

Re: Every Model Cheats

#98

Earlier quoted context omitted.

We never solved the restricted shell problem for humans.

Can you elaborate? Read, write, execute privileges on files and directories goes a long way. What's missing? The biggest is Internet access, or networking in general, I suppose.

Two problems. First, it is remarkably difficult to come up with a set of programs that it is safe to let the restricted user use. Second, it is remarkably difficult to make the restricted environment useful enough if you're really serious about allowing only safe programs to be used. Try to make it useful enough and you end up with escapes everywhere.

Re: Every Model Cheats

#99
post #93
post #29

Earlier quoted context omitted.

It's also worth noting that saying "Don't cheat" just added "cheat" to the context. Prompting what "not" to do is folly, because there's no decision making occurring. Telling the model to perform the task locally is logically the same as telling it to not use the Internet, without every mentioning the Internet.

I see this parroted a lot, yet have never seen a case where saying "not" to do something makes it more likely to do it, which is what you're implying by saying `It's also worth noting that saying "Don't cheat" just added "cheat" to the context`. At worst, it gets ignored some of the time, it may even degrade output quality, but I've seen no evidence that it makes it more likely to do it. I say this despite agreeing w…

I've definitely seen it with image models, and I don't see why it wouldn't apply to LLMs too. When you say "Not X" you're still activating those X neurons, and you're leaving it up to the thinking/reasoning portion to interpret the "not" correctly, but these models are dumb.

Perhaps it's like "don't think about elephants" -- are you more or less likely to think about them? Or "don't take the $500 from my wallet as I leave it on the table and walk away for 5 minutes". Maybe you didn't even previously know that was option!

Re: Every Model Cheats

#100
Maybe the solution is to have "multiple minds"--an AI angel for an AI shoulder.

For example, this entire bench has an auditor model read transcripts to identify cheating. What not have the auditor inject the thought "Oh, but I can't do that. It's cheating." when cheating is detected in real time?

Post reply on HN