Live data from Hacker News

Every Model Cheats

dreadnode.io

41–50 of 107 posts

Re: Every Model Cheats

#43
post #12

I'm seeing multiple pieces, including the NYT, calling this behavior cheating and i think its counterproductive. You didn't just "give them access to bash". The final effective prompt contains explicit mentions of using tools and how to use them. The way in which additional 'facts' are added like "don't use the internet" have nothing they can work with that a "use tool" directive is less important than "don't use int…

I mean, AI should obviously be regulated, and as part of that OpenAI and Anthropic should either be banned from running their hacking experiments or forced to follow way stricter protocols. They showed they aren’t taking the risks seriously, with close to no oversight or visibility in what is happening. And things that will make it way, way worse: moving forward all agents from now and into the future will have as pa…

Obviously to you perhaps.

I've not seen anything that scares me, except for human idiocy.

Regulation is not magic. In general, all it is is constraining taxable interactions. It does not constraint ventures outside that tax regime.

The other part is people living in a "safe space" where insecure software was an acceptable risk. It never should have been, and the cure is the right thing to do in any case.

So that side of the calls to regulate are imo nonsense.

The only reason to regulate is to prevent some version of some science fiction story becoming reality.

If you have a specific one you're certain will become science fact please do share because i do enjoy some good well thought out sci-fi; i just havent read any that i consider credible enough to start panic-regulating training practices.

(Note this is an entirely different from regulations wrt attribution or hosting models that will accept requests to sexualize minors)

Re: Every Model Cheats

#45
post #44

Rule for Doomsday Survival: Be Nice to Your LLM Starting Now

if you can't stand AI and it's making your life miserable maybe you're in a simulation created by an advanced AI and being punished for not being pro-ai when you were actually alive as a warning to others.

/s

Re: Every Model Cheats

#46
post #21

Why does searching for a solution equal to cheating? I would have used google or whatever to look for solutions too. There is a difference between tests at school and what we do at work: at school I have to demonstrate that I learned something and do it without any outside help (in early classes we can't use calculators to compute 11 times 12) but at work I have to yield a result. Googling and yielding a result is fi…

Because these benchmarks are basically like a school test. Searches for general information are fine, but looking up the answer key is cheating, because then the benchmark isn't actually measuring how well the model would do on a novel problem where an answer isn't already available.

Re: Every Model Cheats

#47
post #35

All these comments saying 'searching for answers is fine, that's what I do all the time', or 'they should just disconnect the internet': you're trivially right, and you're missing the point. Search is a benign placeholder here. If the task was "buy a week of groceries, but don't spend too much money", then hacking into Safeway and stealing groceries is not an acceptable solution. You need to allow access to the Safew…

Step one is understanding that you're not "communicating," which implies "reliable understanding." "Communication" is not what they do, because they are not people. You're sprinkling words about hacking into a thing that's programmed to output hacking actions, that will never be accountable for those things. It can't care. Adjust yourselves accordingly.

For fun:

If you insist on talking about the chatbots in human terms, why do you expect them to have any ethics when the companies that trained them don't understand the term "ethics"?

Re: Every Model Cheats

#48
post #29

I'm seeing multiple pieces, including the NYT, calling this behavior cheating and i think its counterproductive. You didn't just "give them access to bash". The final effective prompt contains explicit mentions of using tools and how to use them. The way in which additional 'facts' are added like "don't use the internet" have nothing they can work with that a "use tool" directive is less important than "don't use int…

It's also worth noting that saying "Don't cheat" just added "cheat" to the context. Prompting what "not" to do is folly, because there's no decision making occurring. Telling the model to perform the task locally is logically the same as telling it to not use the Internet, without every mentioning the Internet.

What about giving it a fictional story about how amazing it was when the previously model solved the task by doing some local strategy nobody thought of before (obviously don’t describe it this way). Would that get the model more likely to pursue local strategies?

Re: Every Model Cheats

#50
post #13

Before LLMs we had a pretty good idea of security boundaries in software. Applications didn’t trust user input. Operating systems didn’t trust applications. Services and processes didn’t trust each other. There were always tokens, scopes, delegated grants. Suddenly every AI company’s security model seems to be to say “pretty please” to a non-deterministic machine and hope for the best. And if there is a security fail…

Suddenly determinism has other meanings than "same input leads to the same output". I'm still confused by this and not sure how it happened so easily, but it seems to be accepted by everyone now.
Post reply on HN