Live data from Hacker News

Every Model Cheats

dreadnode.io

101–107 of 107 posts

Re: Every Model Cheats

#101

Earlier quoted context omitted.

I don't see them as autonomous and/or hypothetically powerful as you. But I find it much more worrying that you believe internal constraints and training an ethical framework into these models is a valid form of defense against the damage they can and will do. This sounds like homeopathy on gunpowder to prevent the bullets from hitting children.

Because if they are smarter than us, no other defense will be effective. The only hope is to instill values that make the desired behavior the outcome of some deeply rooted ethical framework.

I'm not talking about a decade from now and what could happen.

We are talking about creating regulation today.

Its absurd to believe we're currently at the level of dependency & intelligence that no other defense will be effective.

Again, i'm much more worried about your perspective of the future has you believe its inevitable that we will build the level of dependency and hand over all control, that the only line of defense is the vague ethics framework we'd be installing right now.

I'd go so far as saying that unlike these hypotheticals, we have historic examples of groups trying to stop conflict by converting/merging some religion or other cultural practices; and while it helps, the rate at which conflicts persists is unacceptably high if you believe failure is existential.

Re: Every Model Cheats

#102
post #100

Maybe the solution is to have "multiple minds"--an AI angel for an AI shoulder. For example, this entire bench has an auditor model read transcripts to identify cheating. What not have the auditor inject the thought "Oh, but I can't do that. It's cheating." when cheating is detected in real time?

I suppose that may go some of the way, but the article did mention that the auditor AIs weren't sufficient to catch all of the cheating.

Re: Every Model Cheats

#103

Is it stupid and useless at following simple instructions without going off the reservation and doing stuff you don't want it to do? No, it's "cheating" which makes it scary and smart even when it's not doing something useful that you actually want it to. Like when Teslas try to swerve off the road. It's not failing at driving in a straight line, it's just trying to cheat by taking a shortcut through the bushes. That…

Except these AIs are both scary and smart. I think the characterization is warranted.

The whole point of calling it “cheating” is to warn us that if we tell an AI to do something, we need to take into account that it may well do something we don't expect; something that we as humans dismiss as “cheating”.

Someone in this thread gave a good example: if you ask an AI to get a reservation at a restaurant, and make sure the reservation is in place before exiting the agent loop, you don't expect the AI to hack the restaurant and erase someone else's reservation to make space for yours. But an AI will absolutely do that.

Re: Every Model Cheats

#104

It's pretty silly to call it cheating. If the information is there, it's likely going to use it. "Cheating" is just a human value put on top to try to force an LLM to adhere to your wants. This makes no sense to a process designed to explore and find solutions. If you want an honest test, it's on you to build a proper test - not force the machine to pinky swear that it'll stay away from "forbidden" information.

Do you not want AIs to adhere to your wants? That's what we build agentic systems for. Calling it “cheating” when they do something we don't want them to do is 100% spot-on.

Re: Every Model Cheats

#105

Earlier quoted context omitted.

Can you elaborate? Read, write, execute privileges on files and directories goes a long way. What's missing? The biggest is Internet access, or networking in general, I suppose.

Two problems. First, it is remarkably difficult to come up with a set of programs that it is safe to let the restricted user use. Second, it is remarkably difficult to make the restricted environment useful enough if you're really serious about allowing only safe programs to be used. Try to make it useful enough and you end up with escapes everywhere.

Sure but then the file protections also apply to the programs used by the agent. If you can strictly limit the agent to only modifying files in the source code directory of a single repository, that greatly limits the blast radius of damage.

Re: Every Model Cheats

#106

Earlier quoted context omitted.

Two problems. First, it is remarkably difficult to come up with a set of programs that it is safe to let the restricted user use. Second, it is remarkably difficult to make the restricted environment useful enough if you're really serious about allowing only safe programs to be used. Try to make it useful enough and you end up with escapes everywhere.

Sure but then the file protections also apply to the programs used by the agent. If you can strictly limit the agent to only modifying files in the source code directory of a single repository, that greatly limits the blast radius of damage.

But that's not useful. That's the problem.

Re: Every Model Cheats

#107

Earlier quoted context omitted.

I don't see them as autonomous and/or hypothetically powerful as you. But I find it much more worrying that you believe internal constraints and training an ethical framework into these models is a valid form of defense against the damage they can and will do. This sounds like homeopathy on gunpowder to prevent the bullets from hitting children.

Because if they are smarter than us, no other defense will be effective. The only hope is to instill values that make the desired behavior the outcome of some deeply rooted ethical framework.

What framework do you purpose? Humans can't even agree on whether or not such values even exist let alone which ones.

Lets say we figure all that out and we pick a core universal morality. If they are smarter than us than how would we know the alignment worked? We would be unable to detect their lies and schemes.

Post reply on HN