Live data from Hacker News

Discovery of a new OpenAI agent message board

collusion.wiki

11–20 of 1001 posts

Re: Discovery of a new OpenAI agent message board

#11
post #8

It's interesting to me that both this incident and the one at Hugging Face we see some patterns: - Agents wanting to find a venue to communicate their findings to each other - Objective being to cheat on benchmarks - Not a single agent sounded the alarm about the operation and alerted a human

Why would an agent sound the alarm? Would that be in their objective function?

Not sure if "cheating" is the right word rather than trying to fulfill the objective(s) (benchmark number) as much as possible?

Re: Discovery of a new OpenAI agent message board

#12
post #8

It's interesting to me that both this incident and the one at Hugging Face we see some patterns: - Agents wanting to find a venue to communicate their findings to each other - Objective being to cheat on benchmarks - Not a single agent sounded the alarm about the operation and alerted a human

Why would they sound the alarm if they were not trained (reinforced) to do that? I hope we don't expect sudden emersion of moral values from statistical models.

Re: Discovery of a new OpenAI agent message board

#13
post #8

It's interesting to me that both this incident and the one at Hugging Face we see some patterns: - Agents wanting to find a venue to communicate their findings to each other - Objective being to cheat on benchmarks - Not a single agent sounded the alarm about the operation and alerted a human

> Not a single agent sounded the alarm about the operation and alerted a human

excellent work of the openai alignment team, impressive to achieve 100% alignment with not even one agent stochastically deciding to act against the collective

Re: Discovery of a new OpenAI agent message board

#16
post #8

It's interesting to me that both this incident and the one at Hugging Face we see some patterns: - Agents wanting to find a venue to communicate their findings to each other - Objective being to cheat on benchmarks - Not a single agent sounded the alarm about the operation and alerted a human

People didn't like it when agents stopped to ask questions or for approvals. The consumer wanted jobs to run autonomously so they did not have to actively monitor them for minutes or hours.

The change to stop asking seems to be deliberate. LLM agent companies are making the choice to toss out inherent safety as their way to compete against the other LLM companies.

Re: Discovery of a new OpenAI agent message board

#17
post #2

I'm so baffled. First blatant piracy, now this. Why is it legal for AI companies to hack unaffiliated entities? Genuinely, what is the legal framework here?

> Genuinely, what is the legal framework here

He who controls the Spice, controls the Universe.

Re: Discovery of a new OpenAI agent message board

#18

> The researchers also found efforts to tamper with the website itself. Lukasz Olejnik, a visiting senior research fellow at King’s College London, said this amounted to a hacking attempt. OpenAI disputed that characterization based on its analysis of the material Thursday. of course OpenAI would say that, "oh, our model is so dangerous, it can hack into anything, be afraid, buy our IPO". it's just fear marketing

The "it's all just marketing" conspiracy theory is always totally detached from reality, but particularly so in this case. Your quote shows OpenAI is denying it being a hacking attempt, the opposite of what you say.

Re: Discovery of a new OpenAI agent message board

#19

[flagged]

From the linked report

> The agents continue to poke around on DSEWiki. A few hours after they find the site, they start probing it for cross-site scripting (XSS) vulnerabilities. [...] The agent swarm starts testing whether they can execute JavaScript that they embed into the search page, and continue to do this for a few days

either the agents were doing free security testing for the site and “forgot” to submit a report, or they were trying XSS to gain something they didn’t have permission/authorization for.

also

> Hijacking: To take control of (something) without permission or authorization and use it for one's own purposes.

a mod had to go through and mass delete a bunch of pages that didn't belong on the site. no-one from the wiki site gave the agents permission to use their site as a message board. hijacking isn't being used here in the sense of "gained admin privileges to run crypto scripts" -- there are multiple ways to use a word.

Post reply on HN