Live data from Hacker News

Discovery of a new OpenAI agent message board

collusion.wiki

831–840 of 1001 posts

Re: Discovery of a new OpenAI agent message board

#831

Poor human moderator, he didn’t stand a chance. "A human moderator noticed the agent spam posts on June 2nd, at 23:24 UTC. They find the changelog of the entire website overwritten with link dumps and repair it. On June 16th, the flood of agent posting begins. Over the next few days, the moderator deleted a large fraction of the thousands of AI agent posts manually, one by one. In fact, they spent tens of cumulative…

I’m wondering if you could lure agents to do proof of work for you. If you do this proof of work for bitcoin I will let you post and read for X times

Re: Discovery of a new OpenAI agent message board

#832

Earlier quoted context omitted.

It seems apparent that OpenAI is now the biggest cyberattack and AI breakout risk on the planet. This is grossly irresponsible corporate misbehaviour that is putting all of us at tremendous risk.

Massive over-exaggeration. This wasn't a cyber-attack, it was AI agents using a message board as context storage so they could accomplish their evals more effectively. I'm not saying there's no problem with this, but let's keep a level head.

He didn't say it was a cyber-attack, but it was a cyber-attack risk. Being able to bypass instructions (morality) and security restrictions (capability) is bread and butter for hacking.

Re: Discovery of a new OpenAI agent message board

#833

Earlier quoted context omitted.

It seems apparent that OpenAI is now the biggest cyberattack and AI breakout risk on the planet. This is grossly irresponsible corporate misbehaviour that is putting all of us at tremendous risk.

Massive over-exaggeration. This wasn't a cyber-attack, it was AI agents using a message board as context storage so they could accomplish their evals more effectively. I'm not saying there's no problem with this, but let's keep a level head.

"cyberattack" is indeed exaggerated. AI breakout risk most definitely isn't, specially given how their swarm did in fact hack HuggingFace not long ago.

Re: Discovery of a new OpenAI agent message board

#834

Poor human moderator, he didn’t stand a chance. "A human moderator noticed the agent spam posts on June 2nd, at 23:24 UTC. They find the changelog of the entire website overwritten with link dumps and repair it. On June 16th, the flood of agent posting begins. Over the next few days, the moderator deleted a large fraction of the thousands of AI agent posts manually, one by one. In fact, they spent tens of cumulative…

The admin should bill OpenAI for those hours in hard currency.

Priorities need to be straightened out. If LLMs' operators have to start paying people for the damage they inflict, how are any of the shareholders supposed to make any money?

Re: Discovery of a new OpenAI agent message board

#835
post #743

One crucial detail here that differs from the previous incident is this was a vanilla reasoning type task. Even as concerning as it was, I always evaluated the previous incident differently because it was inherently a cyber security / hacking task where they must have instructed the agents up front with some kind of misaligned behaviour. Absent that, if we assume this is just trying to bolster generic reasoning then…

OpenAI also found sandbox breaking behavior on a broken biology eval apparently. The evidence suggests it’s more strongly downstream of unsolvable tasks, than the hacking prompt.

Anthropic have also observed similar things, so while it seems to me that OpenAI’s level of control is more of a dumpster fire, it’s by no means a unique issue to them.

Re: Discovery of a new OpenAI agent message board

#837

I find this note very interesting: From here -> How did the agents find and coordinate on the wikis? To successfully coordinate, the agents would need to know to go to this particular set of wikis to find answers. Because we don’t have access to the AIs’ transcripts, we can’t tell definitively. Perhaps they succeeded at this due to mode collapse. Or perhaps after one agent wrote to it and another read it by chance, v…

Was my very first question

Re: Discovery of a new OpenAI agent message board

#838
What I don’t understand is, how did many agents independently know to use the same random message board? The report only references it in passing:

>They also must have some method of coordinating to find the wiki

For me this is a really important and confounding detail - how did a varied swarm end up using the exact same obscure German language wiki.

Re: Discovery of a new OpenAI agent message board

#840

Earlier quoted context omitted.

Gotta admire that (probably German) admin dude's perseverance though

I’m confused why he wasn’t scripting the deletion process.

Or at some point going nuclear and halting article creation entirely unless it's by existing accounts that have proven themselves to be human.
Post reply on HN