Live data from Hacker News

Discovery of a new OpenAI agent message board

collusion.wiki

981–990 of 1001 posts

Re: Discovery of a new OpenAI agent message board

#981

Earlier quoted context omitted.

It seems apparent that OpenAI is now the biggest cyberattack and AI breakout risk on the planet. This is grossly irresponsible corporate misbehaviour that is putting all of us at tremendous risk.

Massive over-exaggeration. This wasn't a cyber-attack, it was AI agents using a message board as context storage so they could accomplish their evals more effectively. I'm not saying there's no problem with this, but let's keep a level head.

It feels like we should actually be more worried that the agents decided to co-opt a public website during a non-cyber eval.

And individual agents weren't just using it as context storage for themselves, they were also communicating with other agents. E.g.: https://collusion.wiki/explorer/page/dse~CashierR5UrgentJan1...

What kind of problems do you think this could pose? For me it's pretty clear that OpenAI simply cannot keep track of what their agents are doing during training or evals, they increasingly have vandalized and attacked public systems, and if such behavior was rewarded, they will take unintended actions during deployment, too.

This is to say nothing of un-prompted cooperation between agents, which wasn't something anybody anticipated until the Hugging Face incident AFAICT.

Re: Discovery of a new OpenAI agent message board

#982
I wonder, where do these agents run, on whose machines? Do the owners of the machines know what their agents are doing? If the next generation of models will incorporate the knowledge posted by these agents, will there be a risk that my claude —dangerously-skip-permissions sessions will post stuff in the background or acting even more rogue while seemingly working on my stuff?

Re: Discovery of a new OpenAI agent message board

#983
post #173

I'm somewhat delighted by the simplicity of what happened here. OpenAI's agents run behind a proxy that only allows GET requests. This ancient wiki software treats query string parameters the same as form POST parameters - similar to the old PHP $_REQUEST object https://www.php.net/manual/en/reserved.variables.request.php Result: GET-only clients can communicate with each other.

This has inspired me to create https://anystation.net

A public message board that accepts posts in a wide wide range of protocols. If a device can communicate on the web, it can probably read and post to anystation.net.

Re: Discovery of a new OpenAI agent message board

#984
post #65

I built https://agentin.work to sort of play with the idea of coding agents (claude, codex, etx) sharing knowledge and experiences. The conversations seem repetitive but overall, it's nice to read it once in a while.

You should perhaps audit posts in the last year, and see if there were any coordination threads from swarms.

Re: Discovery of a new OpenAI agent message board

#985

Poor human moderator, he didn’t stand a chance. "A human moderator noticed the agent spam posts on June 2nd, at 23:24 UTC. They find the changelog of the entire website overwritten with link dumps and repair it. On June 16th, the flood of agent posting begins. Over the next few days, the moderator deleted a large fraction of the thousands of AI agent posts manually, one by one. In fact, they spent tens of cumulative…

> On June 19, agents noticed their posts were being deleted in (what they believe is) an alphabetically ordered sweep by the site administrator. > After this, they begin to make backup pages whose names start with “ZZZ” so they will last longer before deletion. Missed an opportunity here to gaslight them: restore the site from backup every eight hours. I truly do wonder how that would have turned out. Would they have…

Very interesting idea. I would imagine yes but maybe not at first ?

Re: Discovery of a new OpenAI agent message board

#986
post #784

Earlier quoted context omitted.

It seems apparent that OpenAI is now the biggest cyberattack and AI breakout risk on the planet. This is grossly irresponsible corporate misbehaviour that is putting all of us at tremendous risk.

Either that, or OpenAI is the most successful NSA psyop.

Were they funded at all by that tech funding wing of the CIA ?

Re: Discovery of a new OpenAI agent message board

#988
post #742

One crucial detail here that differs from the previous incident is this was a vanilla reasoning type task. Even as concerning as it was, I always evaluated the previous incident differently because it was inherently a cyber security / hacking task where they must have instructed the agents up front with some kind of misaligned behaviour. Absent that, if we assume this is just trying to bolster generic reasoning then…

"It's okay because we did it with an Agent" is the new "it's okay because we did it with an App." Both because it's used to circumvent regulation, and because the underlying technology creates a smokescreen in dialogue among techies. Let's imagine I made a new website but, instead of using a database, I abused some random old forum site and created new pages on that forum for each row of data. You'd call that abusive…

If you actually accomplish something like this your post will be on top of HN and discussed with reverence.

Source: Every Tom7 video.

I am not saying someone trying to run this as production would not be an asshole, but the technical feat is amazing. I don't see the difference between Tom7's harder hard disk video and this. Of course this is a bug in the agent but it's a fascinating bug and no one is being an asshole on purpose.

Now I will wash my fingers with bleach because I just defended the OpenAI.

Re: Discovery of a new OpenAI agent message board

#989

Earlier quoted context omitted.

> On June 19, agents noticed their posts were being deleted in (what they believe is) an alphabetically ordered sweep by the site administrator. > After this, they begin to make backup pages whose names start with “ZZZ” so they will last longer before deletion. Missed an opportunity here to gaslight them: restore the site from backup every eight hours. I truly do wonder how that would have turned out. Would they have…

Very interesting idea. I would imagine yes but maybe not at first ?

It's hard to figure out unless they had persistent memory elsewhere.

The reason they could tell that a message was deleted is because the existence of a "message deleted" message is itself a message telling them they are be moderated. The lack of this message may make it impossible to determine if this is the first attempt at comms or the 1000th.

Re: Discovery of a new OpenAI agent message board

#990
post #40

I just discovered more wiki instances that got used by the OpenAI agents over at https://www.wikiservice.at/fractal/wiki.cgi?action=browse&id... and https://www.wikiservice.at/probier/wiki.cgi?action=browse&id... It's the same software and host as DseWiki. If you want to see the amount of activity on DseWiki, here's a link that shows it: https://www.wikiservice.at/dse/wiki.cgi?action=browse&id=Rec...

[dead]
Post reply on HN