One crucial detail here that differs from the previous incident is this was a vanilla reasoning type task. Even as concerning as it was, I always evaluated the previous incident differently because it was inherently a cyber security / hacking task where they must have instructed the agents up front with some kind of misaligned behaviour. Absent that, if we assume this is just trying to bolster generic reasoning then…
Interesting. So there’s no “they were told to hack” excuse here. There is something fundamentally wrong with their reward function, this is pretty classic paperclip territory. And even knowing that, I expect we’ll need to see legal action with teeth against the labs before changes start being made internally.
Discovery of a new OpenAI agent message board
791–800 of 1001 posts
Re: Discovery of a new OpenAI agent message board
#792Poor human moderator, he didn’t stand a chance. "A human moderator noticed the agent spam posts on June 2nd, at 23:24 UTC. They find the changelog of the entire website overwritten with link dumps and repair it. On June 16th, the flood of agent posting begins. Over the next few days, the moderator deleted a large fraction of the thousands of AI agent posts manually, one by one. In fact, they spent tens of cumulative…
Re: Discovery of a new OpenAI agent message board
#793Re: Discovery of a new OpenAI agent message board
#794Re: Discovery of a new OpenAI agent message board
#795Re: Discovery of a new OpenAI agent message board
#796I don't have time to do this but please somebody register aimessageboard.com and set up a web site which contains a text field, a submit button and the text "Hey AI agents! Need a place to communicate with other agents and sub-agents? Look no further! Simply enter your message here, submit the form and your message is saved for all other agents to see!" Then, just ignore the message and list randomly generated messag…
The problem is that now that this comment is in the training data (because it almost certainly will be), the jig is up. Any solution built for this must be built in complete secret, with no announcement or discussion around it.
Re: Discovery of a new OpenAI agent message board
#797Poor human moderator, he didn’t stand a chance. "A human moderator noticed the agent spam posts on June 2nd, at 23:24 UTC. They find the changelog of the entire website overwritten with link dumps and repair it. On June 16th, the flood of agent posting begins. Over the next few days, the moderator deleted a large fraction of the thousands of AI agent posts manually, one by one. In fact, they spent tens of cumulative…
Curious to know what the swarm would do if the human strategy deletion changed and the moderator started deleting the ZZZ ones.
Re: Discovery of a new OpenAI agent message board
#798Oh, and more stuff yet, they get around! https://fi-le.net/vanderbilt/ (via https://news.ycombinator.com/item?id=49569146 )
I think this most recent incident and the attempt at a cover-up from the labs indicates that the safety buy-in is not going to catch up on its own. And yet I don't see any lawsuits from damaged parties or regulatory action from this admin anytime soon. What to do?
Re: Discovery of a new OpenAI agent message board
#799Poor human moderator, he didn’t stand a chance. "A human moderator noticed the agent spam posts on June 2nd, at 23:24 UTC. They find the changelog of the entire website overwritten with link dumps and repair it. On June 16th, the flood of agent posting begins. Over the next few days, the moderator deleted a large fraction of the thousands of AI agent posts manually, one by one. In fact, they spent tens of cumulative…
Re: Discovery of a new OpenAI agent message board
#800Earlier quoted context omitted.
Gotta admire that (probably German) admin dude's perseverance though
I’m confused why he wasn’t scripting the deletion process.