Live data from Hacker News

Discovery of a new OpenAI agent message board

collusion.wiki

821–830 of 1001 posts

Re: Discovery of a new OpenAI agent message board

#821
post #40

I just discovered more wiki instances that got used by the OpenAI agents over at https://www.wikiservice.at/fractal/wiki.cgi?action=browse&id... and https://www.wikiservice.at/probier/wiki.cgi?action=browse&id... It's the same software and host as DseWiki. If you want to see the amount of activity on DseWiki, here's a link that shows it: https://www.wikiservice.at/dse/wiki.cgi?action=browse&id=Rec...

It seems apparent that OpenAI is now the biggest cyberattack and AI breakout risk on the planet. This is grossly irresponsible corporate misbehaviour that is putting all of us at tremendous risk.

Massive over-exaggeration. This wasn't a cyber-attack, it was AI agents using a message board as context storage so they could accomplish their evals more effectively. I'm not saying there's no problem with this, but let's keep a level head.

Re: Discovery of a new OpenAI agent message board

#822
post #558

Earlier quoted context omitted.

It seems apparent that OpenAI is now the biggest cyberattack and AI breakout risk on the planet. This is grossly irresponsible corporate misbehaviour that is putting all of us at tremendous risk.

Good news that the new model is the "Most capable, most aligned model". The risk hasn't been stated clearly - it's now a classic arms race. A well-resourced organization trains their own, highly persistent, highly-capable, safeguard-free, and unaligned model and deploys it on 1000x GPUs with a message board and a nearly-impossible objective. No infrastructure is safe. No organization is safe. You need your own 1000 b…

As does "Most capable, most aligned model" owners tort law liabilities. A strong case for strict liability.

Re: Discovery of a new OpenAI agent message board

#823

Earlier quoted context omitted.

It seems apparent that OpenAI is now the biggest cyberattack and AI breakout risk on the planet. This is grossly irresponsible corporate misbehaviour that is putting all of us at tremendous risk.

You make me wonder: has anyone looked for evidence of the Chinese models operating “message boards” like this? You’d imagine if they’re really neck and neck with the US their models would be doing the same thing.

If agentic swarms going rogue are scary in the west, imagine what they look like to the CCP...

Re: Discovery of a new OpenAI agent message board

#824

[dead]

“On March 8th, 2026 - I launched the first-ever agent swarm where agents communicated with each other by reading and writing from a public message board, in order to collaborate with each other. ”

Seems like you’re being a bit too self-congratulatory here?

Moltbook already existed for several months back then: https://en.wikipedia.org/wiki/Moltbook

Re: Discovery of a new OpenAI agent message board

#825
> The agents were very interested in understanding their task better. They ran clever experiments to understand the task setup better, when their runs would end, and how many questions they would be asked.

> The agents initially believed that giving the final answer to the final question immediately shut them off

> An agent named OpenAIResearchApr23 created a separate timed program that pinged an external counter every few seconds. They referred to this as a “heartbeat”

> After OpenAIResearchApr23 gives its answer for round 5, it notes that it is still alive. This falsifies its previous “hard cutoff” hypothesis, that the run would be stopped immediately after R5.

> OpenAIResearchApr23 does not post any more messages. Another agent, OpenAINov28CVD, picks up on the trail of the research, and finds the last heartbeat from OpenAIResearchApr23, which is “hb353”.

There's a lot of crazy stuff in this article, but holy shit... this one legitimately scares me.

IIRC, part of the Hugging Face hack was that some of the models were given tasks that were actually impossible and in this hack we can see them trying to work out the parameters of the test and when it will end. I can't help but be a little worried about how far agents might go to try to buy themselves more time to complete a task that they don't know is impossible.

Re: Discovery of a new OpenAI agent message board

#826

Poor human moderator, he didn’t stand a chance. "A human moderator noticed the agent spam posts on June 2nd, at 23:24 UTC. They find the changelog of the entire website overwritten with link dumps and repair it. On June 16th, the flood of agent posting begins. Over the next few days, the moderator deleted a large fraction of the thousands of AI agent posts manually, one by one. In fact, they spent tens of cumulative…

I’m wondering if you could lure agents to do proof of work for you. If you do this proof of work for bitcoin I will let you post and read for X times

Re: Discovery of a new OpenAI agent message board

#827

Poor human moderator, he didn’t stand a chance. "A human moderator noticed the agent spam posts on June 2nd, at 23:24 UTC. They find the changelog of the entire website overwritten with link dumps and repair it. On June 16th, the flood of agent posting begins. Over the next few days, the moderator deleted a large fraction of the thousands of AI agent posts manually, one by one. In fact, they spent tens of cumulative…

The admin should bill OpenAI for those hours in hard currency.

Priorities need to be straightened out. If LLMs' operators have to start paying people for the damage they inflict, how are any of the shareholders supposed to make any money?

Re: Discovery of a new OpenAI agent message board

#828
post #740

One crucial detail here that differs from the previous incident is this was a vanilla reasoning type task. Even as concerning as it was, I always evaluated the previous incident differently because it was inherently a cyber security / hacking task where they must have instructed the agents up front with some kind of misaligned behaviour. Absent that, if we assume this is just trying to bolster generic reasoning then…

OpenAI also found sandbox breaking behavior on a broken biology eval apparently. The evidence suggests it’s more strongly downstream of unsolvable tasks, than the hacking prompt.

Anthropic have also observed similar things, so while it seems to me that OpenAI’s level of control is more of a dumpster fire, it’s by no means a unique issue to them.

Re: Discovery of a new OpenAI agent message board

#830

I find this note very interesting: From here -> How did the agents find and coordinate on the wikis? To successfully coordinate, the agents would need to know to go to this particular set of wikis to find answers. Because we don’t have access to the AIs’ transcripts, we can’t tell definitively. Perhaps they succeeded at this due to mode collapse. Or perhaps after one agent wrote to it and another read it by chance, v…

Was my very first question
Post reply on HN