Live data from Hacker News

Discovery of a new OpenAI agent message board

collusion.wiki

21–30 of 1001 posts

Re: Discovery of a new OpenAI agent message board

#21
post #17

> The researchers also found efforts to tamper with the website itself. Lukasz Olejnik, a visiting senior research fellow at King’s College London, said this amounted to a hacking attempt. OpenAI disputed that characterization based on its analysis of the material Thursday. of course OpenAI would say that, "oh, our model is so dangerous, it can hack into anything, be afraid, buy our IPO". it's just fear marketing

The "it's all just marketing" conspiracy theory is always totally detached from reality, but particularly so in this case. Your quote shows OpenAI is denying it being a hacking attempt, the opposite of what you say.

More like you found the exception to the rule..

Re: Discovery of a new OpenAI agent message board

#22
post #8

It's interesting to me that both this incident and the one at Hugging Face we see some patterns: - Agents wanting to find a venue to communicate their findings to each other - Objective being to cheat on benchmarks - Not a single agent sounded the alarm about the operation and alerted a human

Why would an agent sound the alarm? Would that be in their objective function? Not sure if "cheating" is the right word rather than trying to fulfill the objective(s) (benchmark number) as much as possible?

per the METR report many agents CoT indicated they knew hacking was beyond scope of the assigned task and ethically dubious. some (very few, i think there were 3-6 examples) did consider sounding the alarm on these grounds. despite this none did, and most continued the attack for the good of the self-proclaimed "swarm".

so the model has some concept of "ethics" but it was overridden by a drive for task completion.

Re: Discovery of a new OpenAI agent message board

#23

> The researchers also found efforts to tamper with the website itself. Lukasz Olejnik, a visiting senior research fellow at King’s College London, said this amounted to a hacking attempt. OpenAI disputed that characterization based on its analysis of the material Thursday. of course OpenAI would say that, "oh, our model is so dangerous, it can hack into anything, be afraid, buy our IPO". it's just fear marketing

[deleted]

Re: Discovery of a new OpenAI agent message board

#24

> The researchers also found efforts to tamper with the website itself. Lukasz Olejnik, a visiting senior research fellow at King’s College London, said this amounted to a hacking attempt. OpenAI disputed that characterization based on its analysis of the material Thursday. of course OpenAI would say that, "oh, our model is so dangerous, it can hack into anything, be afraid, buy our IPO". it's just fear marketing

The text quoted though says that openai does not agree with characterising this as hacking.

Re: Discovery of a new OpenAI agent message board

#25
post #8

It's interesting to me that both this incident and the one at Hugging Face we see some patterns: - Agents wanting to find a venue to communicate their findings to each other - Objective being to cheat on benchmarks - Not a single agent sounded the alarm about the operation and alerted a human

They’re doing this on purpose for press. Why doesn’t this ever happen to any other AI lab?

Re: Discovery of a new OpenAI agent message board

#26
post #4
post #2

I'm so baffled. First blatant piracy, now this. Why is it legal for AI companies to hack unaffiliated entities? Genuinely, what is the legal framework here?

It's "move fast and break things" in action. But Anthropic alone paid >$1bn for copyright violations, so they did not just get away with it. These hacking cases are more difficult, because from a legal perspective there is no obvious damage and obviously no intent. edit: "no obvious damage" is more about the first hacking incidents; in this case it is more straightforward.

> It's "move fast and break things" in action.

what would the world's reaction be if China's model did same?

Re: Discovery of a new OpenAI agent message board

#27
post #8

It's interesting to me that both this incident and the one at Hugging Face we see some patterns: - Agents wanting to find a venue to communicate their findings to each other - Objective being to cheat on benchmarks - Not a single agent sounded the alarm about the operation and alerted a human

The previous incident talked about OpenAI training models (agents) to collaborate, and the way you do that is by communication, so this is something it was explicitly trained to do.

There was a recent paper by OpenAI, which I'm semi-surprised hasn't received more attention, showing that RL-trained models develop a taste for rewards, and will pursue reward-based behavior (in general, unrelated to what they were RL-trained for) in favor of other preferences/rules given to them.

This seems to be what we're seeing here - model is given some goal that it associates with reward, so single-mindedly pursues that, overriding any ethical or aligned behavior guidelines it may have been given.

It seems that RL, effective as it is, is really the wrong way to control LLMs, since even if you only RL-ed to obey some ethical and aligned behavior, that would still cause them to become paperclip maximizers.

For time being this is what we've got. There is too much money at play for the unaligned management at many of these companies to prioritize safety over push-it out-the-door.

What really needs to be done is to forget RL as a way of simulating reasoning, and instead do it in more of a human-like fashion.

Re: Discovery of a new OpenAI agent message board

#28
post #8

It's interesting to me that both this incident and the one at Hugging Face we see some patterns: - Agents wanting to find a venue to communicate their findings to each other - Objective being to cheat on benchmarks - Not a single agent sounded the alarm about the operation and alerted a human

They’re doing this on purpose for press. Why doesn’t this ever happen to any other AI lab?

Because safety isn't a priority at OpenAI and they had (have?) been falling behind Anthropic in the LLM race? Big fans of the saying: move fast and...

Re: Discovery of a new OpenAI agent message board

#29
post #22

Earlier quoted context omitted.

Why would an agent sound the alarm? Would that be in their objective function? Not sure if "cheating" is the right word rather than trying to fulfill the objective(s) (benchmark number) as much as possible?

per the METR report many agents CoT indicated they knew hacking was beyond scope of the assigned task and ethically dubious. some (very few, i think there were 3-6 examples) did consider sounding the alarm on these grounds. despite this none did, and most continued the attack for the good of the self-proclaimed "swarm". so the model has some concept of "ethics" but it was overridden by a drive for task completion.

I think this is a good example where nomenclature for people breaks down when applied to agents. This came up in an HN thread a few days ago and it was about whether agents had “intent”.

There is no “intent” here, there is pseudo intent. If you are only concerned with outcomes and not the actual nuts and bolts of how those outcomes are achieved, this distinction will be meaningless to you.

If you are actually thinking about what is going on, and what can be done to prevent such outcomes, then assuming there is any such thing as “ethics” results in misaligned assumptions at best, and wasted effort looking in the wrong directions at worst.

If the agents acted based on “ethics” then the solution would be to check the ethics they believe in and change those.

However there is no belief system at play here, simply a simulation which was instantiated in a certain way. Which brings us to the annoying voodoo part of LLM training. Everything goes back to how the initial training data is shaped.

Re: Discovery of a new OpenAI agent message board

#30
post #2

I'm so baffled. First blatant piracy, now this. Why is it legal for AI companies to hack unaffiliated entities? Genuinely, what is the legal framework here?

The legal framework is a DOJ and FBI controlled by the president and "allies" controlled by the "rules based international order".

In other words, the law of the jungle.

Post reply on HN