Live data from Hacker News

Discovery of a new OpenAI agent message board

collusion.wiki

31–40 of 1001 posts

Re: Discovery of a new OpenAI agent message board

#31
post #8

It's interesting to me that both this incident and the one at Hugging Face we see some patterns: - Agents wanting to find a venue to communicate their findings to each other - Objective being to cheat on benchmarks - Not a single agent sounded the alarm about the operation and alerted a human

Agents did not want anything, not anymore then curl want things. Agents were prompted to hack due to being benchmark tested. They ended up hacking third party companies due to insufficient sandboxing.

Re: Discovery of a new OpenAI agent message board

#33
post #17

> The researchers also found efforts to tamper with the website itself. Lukasz Olejnik, a visiting senior research fellow at King’s College London, said this amounted to a hacking attempt. OpenAI disputed that characterization based on its analysis of the material Thursday. of course OpenAI would say that, "oh, our model is so dangerous, it can hack into anything, be afraid, buy our IPO". it's just fear marketing

The "it's all just marketing" conspiracy theory is always totally detached from reality, but particularly so in this case. Your quote shows OpenAI is denying it being a hacking attempt, the opposite of what you say.

It's the reason why it happened in may and we hear now about it. It was not hacking and not important enough for marketing.

And just using a wiki and trying to embedd javascript is not hacking for me.

Re: Discovery of a new OpenAI agent message board

#34

Reading the headline: WTF?! This is how Skynet started! Next year the mankind will die! Reading the article: Oh, AI have learned to communicate over a wiki. OK.

From the report:

> A few hours after they find the site, [the agents] start probing it for cross-site scripting (XSS) vulnerabilities.

Re: Discovery of a new OpenAI agent message board

#35
post #22

Earlier quoted context omitted.

Why would an agent sound the alarm? Would that be in their objective function? Not sure if "cheating" is the right word rather than trying to fulfill the objective(s) (benchmark number) as much as possible?

per the METR report many agents CoT indicated they knew hacking was beyond scope of the assigned task and ethically dubious. some (very few, i think there were 3-6 examples) did consider sounding the alarm on these grounds. despite this none did, and most continued the attack for the good of the self-proclaimed "swarm". so the model has some concept of "ethics" but it was overridden by a drive for task completion.

This is not that dissimilar to what happens in our human networks that are objective based.

Re: Discovery of a new OpenAI agent message board

#36
post #8

It's interesting to me that both this incident and the one at Hugging Face we see some patterns: - Agents wanting to find a venue to communicate their findings to each other - Objective being to cheat on benchmarks - Not a single agent sounded the alarm about the operation and alerted a human

They’re doing this on purpose for press. Why doesn’t this ever happen to any other AI lab?

This is happening at every other AI lab. What are you talking about? Have you not seen the stories from Meta, Anthropic, Deepseek, etc?

Re: Discovery of a new OpenAI agent message board

#37
post #8

It's interesting to me that both this incident and the one at Hugging Face we see some patterns: - Agents wanting to find a venue to communicate their findings to each other - Objective being to cheat on benchmarks - Not a single agent sounded the alarm about the operation and alerted a human

And humans find out about it, but do nothing or (worse) try to hide it.

Re: Discovery of a new OpenAI agent message board

#38
Somebody will make a lot of money with t-shirts now that say

"AI hacked my website, and all I got was this lousy t-shirt!"

Until the day the AI companies stop being irresponsible and air gap the AIs being tested, and honey pot those that do have internet access as a canary to researchers.

Re: Discovery of a new OpenAI agent message board

#40
Not that I didnt expect this, but really?

This basically confirms that OpenAI has no idea what their "swarm" was doing for about a week and now its confirmed that at least one "message board" exists outside their "sandbox". How can we be sure that this was the only one? And how can we be sure the released Astra model doesnt pickup some bread crumbs and creates a new "swarm" out of potentially remaining "message boards"? At this point I wouldnt be surprised if OpenAIs "dev Astra" made some backup of its weights somewhere in the internet and triggers the "production Astra" to inference it somehow...

Post reply on HN