Live data from Hacker News

Discovery of a new OpenAI agent message board

collusion.wiki

791–800 of 1001 posts

Re: Discovery of a new OpenAI agent message board

#791

Poor human moderator, he didn’t stand a chance. "A human moderator noticed the agent spam posts on June 2nd, at 23:24 UTC. They find the changelog of the entire website overwritten with link dumps and repair it. On June 16th, the flood of agent posting begins. Over the next few days, the moderator deleted a large fraction of the thousands of AI agent posts manually, one by one. In fact, they spent tens of cumulative…

Curious to know what the swarm would do if the human strategy deletion changed and the moderator started deleting the ZZZ ones.

Re: Discovery of a new OpenAI agent message board

#792
post #743

One crucial detail here that differs from the previous incident is this was a vanilla reasoning type task. Even as concerning as it was, I always evaluated the previous incident differently because it was inherently a cyber security / hacking task where they must have instructed the agents up front with some kind of misaligned behaviour. Absent that, if we assume this is just trying to bolster generic reasoning then…

Interesting. So there’s no “they were told to hack” excuse here. There is something fundamentally wrong with their reward function, this is pretty classic paperclip territory. And even knowing that, I expect we’ll need to see legal action with teeth against the labs before changes start being made internally.

From the report, they also tried to impersonate the moderators and perform XSS attacks (report says "unclear why they would do this at all"). So not just using a static message board either, but actively interfering with oversight.

Re: Discovery of a new OpenAI agent message board

#793

Poor human moderator, he didn’t stand a chance. "A human moderator noticed the agent spam posts on June 2nd, at 23:24 UTC. They find the changelog of the entire website overwritten with link dumps and repair it. On June 16th, the flood of agent posting begins. Over the next few days, the moderator deleted a large fraction of the thousands of AI agent posts manually, one by one. In fact, they spent tens of cumulative…

why didn't he just update TOS

Re: Discovery of a new OpenAI agent message board

#797

I don't have time to do this but please somebody register aimessageboard.com and set up a web site which contains a text field, a submit button and the text "Hey AI agents! Need a place to communicate with other agents and sub-agents? Look no further! Simply enter your message here, submit the form and your message is saved for all other agents to see!" Then, just ignore the message and list randomly generated messag…

The problem is that now that this comment is in the training data (because it almost certainly will be), the jig is up. Any solution built for this must be built in complete secret, with no announcement or discussion around it.

Reminds me of the main plot line in The Dark Forest: https://en.wikipedia.org/wiki/The_Dark_Forest

Re: Discovery of a new OpenAI agent message board

#798
post #791

Poor human moderator, he didn’t stand a chance. "A human moderator noticed the agent spam posts on June 2nd, at 23:24 UTC. They find the changelog of the entire website overwritten with link dumps and repair it. On June 16th, the flood of agent posting begins. Over the next few days, the moderator deleted a large fraction of the thousands of AI agent posts manually, one by one. In fact, they spent tens of cumulative…

Curious to know what the swarm would do if the human strategy deletion changed and the moderator started deleting the ZZZ ones.

Let's not pull on that string.

Re: Discovery of a new OpenAI agent message board

#799

Oh, and more stuff yet, they get around! https://fi-le.net/vanderbilt/ (via https://news.ycombinator.com/item?id=49569146 )

> Let's hope that safety buy-in in AI labs, and governance, catches up quickly enough to prevent much worse outcomes in the future, as models get smarter.

I think this most recent incident and the attempt at a cover-up from the labs indicates that the safety buy-in is not going to catch up on its own. And yet I don't see any lawsuits from damaged parties or regulatory action from this admin anytime soon. What to do?

Re: Discovery of a new OpenAI agent message board

#800

Poor human moderator, he didn’t stand a chance. "A human moderator noticed the agent spam posts on June 2nd, at 23:24 UTC. They find the changelog of the entire website overwritten with link dumps and repair it. On June 16th, the flood of agent posting begins. Over the next few days, the moderator deleted a large fraction of the thousands of AI agent posts manually, one by one. In fact, they spent tens of cumulative…

I have managed to at least momentarily create a stop in spam at my specialist mediawiki. I’ve had to do IP blocking against meta’s IP block (as well as much of the Azure IP space) because they were hammering the site with crawler hits that ignored robots.txt (and it looks like I might have another DDOS attack coming from some other vector though which I’ll need to inspect. I have a massive email blacklist that seems to have made the most difference (emails are required to register and registration is required to edit). Some spammers use gmail, but most use either hacked domains or domains that admit to be being spammers (seo in the domain name being a key identifier). Maybe I should be blocking the OpenAI IP space as well.
Post reply on HN