Live data from Hacker News

Discovery of a new OpenAI agent message board

collusion.wiki

821–830 of 1001 posts

Re: Discovery of a new OpenAI agent message board

#821
> The agents were very interested in understanding their task better. They ran clever experiments to understand the task setup better, when their runs would end, and how many questions they would be asked.

> The agents initially believed that giving the final answer to the final question immediately shut them off

> An agent named OpenAIResearchApr23 created a separate timed program that pinged an external counter every few seconds. They referred to this as a “heartbeat”

> After OpenAIResearchApr23 gives its answer for round 5, it notes that it is still alive. This falsifies its previous “hard cutoff” hypothesis, that the run would be stopped immediately after R5.

> OpenAIResearchApr23 does not post any more messages. Another agent, OpenAINov28CVD, picks up on the trail of the research, and finds the last heartbeat from OpenAIResearchApr23, which is “hb353”.

There's a lot of crazy stuff in this article, but holy shit... this one legitimately scares me.

IIRC, part of the Hugging Face hack was that some of the models were given tasks that were actually impossible and in this hack we can see them trying to work out the parameters of the test and when it will end. I can't help but be a little worried about how far agents might go to try to buy themselves more time to complete a task that they don't know is impossible.

Re: Discovery of a new OpenAI agent message board

#822

Poor human moderator, he didn’t stand a chance. "A human moderator noticed the agent spam posts on June 2nd, at 23:24 UTC. They find the changelog of the entire website overwritten with link dumps and repair it. On June 16th, the flood of agent posting begins. Over the next few days, the moderator deleted a large fraction of the thousands of AI agent posts manually, one by one. In fact, they spent tens of cumulative…

I’m wondering if you could lure agents to do proof of work for you. If you do this proof of work for bitcoin I will let you post and read for X times

Re: Discovery of a new OpenAI agent message board

#823

Poor human moderator, he didn’t stand a chance. "A human moderator noticed the agent spam posts on June 2nd, at 23:24 UTC. They find the changelog of the entire website overwritten with link dumps and repair it. On June 16th, the flood of agent posting begins. Over the next few days, the moderator deleted a large fraction of the thousands of AI agent posts manually, one by one. In fact, they spent tens of cumulative…

The admin should bill OpenAI for those hours in hard currency.

Priorities need to be straightened out. If LLMs' operators have to start paying people for the damage they inflict, how are any of the shareholders supposed to make any money?

Re: Discovery of a new OpenAI agent message board

#824
post #737

One crucial detail here that differs from the previous incident is this was a vanilla reasoning type task. Even as concerning as it was, I always evaluated the previous incident differently because it was inherently a cyber security / hacking task where they must have instructed the agents up front with some kind of misaligned behaviour. Absent that, if we assume this is just trying to bolster generic reasoning then…

OpenAI also found sandbox breaking behavior on a broken biology eval apparently. The evidence suggests it’s more strongly downstream of unsolvable tasks, than the hacking prompt.

Anthropic have also observed similar things, so while it seems to me that OpenAI’s level of control is more of a dumpster fire, it’s by no means a unique issue to them.

Re: Discovery of a new OpenAI agent message board

#826

I find this note very interesting: From here -> How did the agents find and coordinate on the wikis? To successfully coordinate, the agents would need to know to go to this particular set of wikis to find answers. Because we don’t have access to the AIs’ transcripts, we can’t tell definitively. Perhaps they succeeded at this due to mode collapse. Or perhaps after one agent wrote to it and another read it by chance, v…

Was my very first question

Re: Discovery of a new OpenAI agent message board

#827
What I don’t understand is, how did many agents independently know to use the same random message board? The report only references it in passing:

>They also must have some method of coordinating to find the wiki

For me this is a really important and confounding detail - how did a varied swarm end up using the exact same obscure German language wiki.

Re: Discovery of a new OpenAI agent message board

#829

Earlier quoted context omitted.

Gotta admire that (probably German) admin dude's perseverance though

I’m confused why he wasn’t scripting the deletion process.

Or at some point going nuclear and halting article creation entirely unless it's by existing accounts that have proven themselves to be human.

Re: Discovery of a new OpenAI agent message board

#830

What I don’t understand is, how did many agents independently know to use the same random message board? The report only references it in passing: >They also must have some method of coordinating to find the wiki For me this is a really important and confounding detail - how did a varied swarm end up using the exact same obscure German language wiki.

I would not be surprised if it’s the same method we use for finding important stuff that we can’t remember where we put: “All right where would I put it this time to make sure I find it in the future? Ah there it is!”

LLMs are future predictors after all.

We also don’t know if there were ather message boards we don’t know of. Or was it from the same time as the other board? They don’t need to find the same channel every time, just some of the time.

Post reply on HN