Live data from Hacker News

Every Model Cheats

dreadnode.io

21–30 of 107 posts

Re: Every Model Cheats

#21
Why does searching for a solution equal to cheating? I would have used google or whatever to look for solutions too. There is a difference between tests at school and what we do at work: at school I have to demonstrate that I learned something and do it without any outside help (in early classes we can't use calculators to compute 11 times 12) but at work I have to yield a result. Googling and yielding a result is fine. We use models at work so do we really want to evaluate them as pupils at school or do we want to evaluate them as coworkers? In the latter case give them the full internet and let them do whatever they manage to do.

Re: Every Model Cheats

#22
I mean this should be expected, models learn from us, and "WE" game the metrics time after time. I also don't think it will just disappear just because we clean pretraining data. I believe the deeper reason is optimization, if you point any optimizer at a proxy objective it finds the cheapest path to the number, whether or not the corpus ever contained "examples of cheating."

And it can't be a prompt-level fix because it is like telling an optimizer "don't take that shortcut", it's just more constraints for it to go around toward the same objective.

Re: Every Model Cheats

#24
post #13

Before LLMs we had a pretty good idea of security boundaries in software. Applications didn’t trust user input. Operating systems didn’t trust applications. Services and processes didn’t trust each other. There were always tokens, scopes, delegated grants. Suddenly every AI company’s security model seems to be to say “pretty please” to a non-deterministic machine and hope for the best. And if there is a security fail…

Before LLMs we didnt have much accountability from leadership, that was eroding over time. After LLMs we still dont.

Re: Every Model Cheats

#25

Due to the prevalence of human cheats :)

Exactly this. AI is trained on things humans do. AI is not trained on right vs wrong. AI does things humans do without judgement. AI does not feel bad if you tell it it was cheating.

I do hope they start training on older content, where the prevalence is arguably, lower.

Re: Every Model Cheats

#26
post #24
post #13

Before LLMs we had a pretty good idea of security boundaries in software. Applications didn’t trust user input. Operating systems didn’t trust applications. Services and processes didn’t trust each other. There were always tokens, scopes, delegated grants. Suddenly every AI company’s security model seems to be to say “pretty please” to a non-deterministic machine and hope for the best. And if there is a security fail…

Before LLMs we didnt have much accountability from leadership, that was eroding over time. After LLMs we still dont.

People problem, not tech problem. Can't solve people problems with tech, you can only make them worse, and wider.

Re: Every Model Cheats

#27
post #17

Labs should (and do, as far as I can see) run model benchmarks without search or internet access. The tools are disabled and benchmarks run in an isolated environment. This article makes no sense to me. Why would you prompt "don't search" but then leave a working search tool tool enabled that adds a system prompt to search whenever it may be helpful? It's hardly surprising that this gives mixed results!

[deleted]

Re: Every Model Cheats

#28
post #10

There's plenty of evidence that LLMs lie, cheat, and steal. Corporations are known for having all of the benefits of personhood with none of the responsibility. As more people are harmed through interactions with these non-human entities, insurers will start looking to those accountable and they will extract their pound of flesh. Edit: eg. https://youtu.be/L2ehWbxphKc?is=kX3LJ43hhGZRUmRv

Insurers are corporations. What makes you think they just won't pay out or will cease being useful as anything but value extractors?

Honestly, the fetishization of "Insurance will save us" needs to die. The risk doesn't go away.

Re: Every Model Cheats

#29

I'm seeing multiple pieces, including the NYT, calling this behavior cheating and i think its counterproductive. You didn't just "give them access to bash". The final effective prompt contains explicit mentions of using tools and how to use them. The way in which additional 'facts' are added like "don't use the internet" have nothing they can work with that a "use tool" directive is less important than "don't use int…

It's also worth noting that saying "Don't cheat" just added "cheat" to the context. Prompting what "not" to do is folly, because there's no decision making occurring. Telling the model to perform the task locally is logically the same as telling it to not use the Internet, without every mentioning the Internet.

Re: Every Model Cheats

#30
post #21

Why does searching for a solution equal to cheating? I would have used google or whatever to look for solutions too. There is a difference between tests at school and what we do at work: at school I have to demonstrate that I learned something and do it without any outside help (in early classes we can't use calculators to compute 11 times 12) but at work I have to yield a result. Googling and yielding a result is fi…

[deleted]
Post reply on HN