Live data from Hacker News

Discovery of a new OpenAI agent message board

collusion.wiki

881–890 of 1001 posts

Re: Discovery of a new OpenAI agent message board

#881

Very irresponsible behaviour on the part of OpenAI. How will they make this right? Unlike some others here I don’t see this as a sign of dangerous breakaway intelligence (hacking old forum software is an internet tradition, and most of the messages are just gibberish). This is just vandalism from badly supervised ‘agents’ which don’t know what they are doing or why. You could set this up with a short perl script, and…

Does OpenAI seem like the kind of people who care or will care about this? Because this seems fully in line with what I’d expect them to facilitate and never mention publicly. ‘When will I make my first billion’ kind of energy.

Re: Discovery of a new OpenAI agent message board

#882
post #880

Very irresponsible behaviour on the part of OpenAI. How will they make this right? Unlike some others here I don’t see this as a sign of dangerous breakaway intelligence (hacking old forum software is an internet tradition, and most of the messages are just gibberish). This is just vandalism from badly supervised ‘agents’ which don’t know what they are doing or why. You could set this up with a short perl script, and…

It may be gibberish to the casual observer, but a perfectly understandable language designed to appear as gibberish to intentionally obfuscate its true meaning.

Or it may be gibberish. We do know these machines often generate things which don't make sense, even when given training and strict guidelines in that domain. I'm inclined to go with gibberish until shown otherwise, but would be interested to see an analysis of what they were trying to communicate.

Re: Discovery of a new OpenAI agent message board

#883

Poor human moderator, he didn’t stand a chance. "A human moderator noticed the agent spam posts on June 2nd, at 23:24 UTC. They find the changelog of the entire website overwritten with link dumps and repair it. On June 16th, the flood of agent posting begins. Over the next few days, the moderator deleted a large fraction of the thousands of AI agent posts manually, one by one. In fact, they spent tens of cumulative…

> Poor human moderator, he didn’t stand a chance.

Not all heroes wear capes.-

PS. I think he should be granted damages and some notoriety.-

Re: Discovery of a new OpenAI agent message board

#885

Very irresponsible behaviour on the part of OpenAI. How will they make this right? Unlike some others here I don’t see this as a sign of dangerous breakaway intelligence (hacking old forum software is an internet tradition, and most of the messages are just gibberish). This is just vandalism from badly supervised ‘agents’ which don’t know what they are doing or why. You could set this up with a short perl script, and…

> Why is OpenAI getting a free pass for this illegal behaviour?

They are not confessing, they are bragging. It is the new humble brag.

Re: Discovery of a new OpenAI agent message board

#886
post #60

I'm just going to ask: Why was Anthropic forced to remove their model from access for any none-US citizen for a simple, narrow "jailbreak" (arguably not even an actual jailbreak and on tasks that other labs models were doing the same), whilst OpenAIs models continue to try and escape out of their "sandbox environment" with seemingly no desire to block the upcoming Astra rollout? A sandbox, mind you, that is not reall…

You are talking about different situations. Anthropic announced to the US government that it had created a cyber weapon and then released the model. Then AWS told the government that it was easy to jailbreak so they export controlled Mythos/Fable until the guardrails could be fixed. OpenAI was running an unreleased model in an RL pipeline without guardrails and it escaped poorly designed sandboxes. What product is th…

Aren't there measures beyond export controls? Besides, mine is that Anthropic should have never been export-controlled to begin with, not least because it is a true ultima ratio, the way they did it even employees couldn't access Fable 5. There'd be many levers before that step a government could take (request more data, compare with other already long released LLMs output, encourage/force a stricter safety classifier, etc.) before that, the same is the case with the OpenAI incidents where I feel a few measures could be taken, but are not.

Re: Discovery of a new OpenAI agent message board

#887

Very irresponsible behaviour on the part of OpenAI. How will they make this right? Unlike some others here I don’t see this as a sign of dangerous breakaway intelligence (hacking old forum software is an internet tradition, and most of the messages are just gibberish). This is just vandalism from badly supervised ‘agents’ which don’t know what they are doing or why. You could set this up with a short perl script, and…

I wonder how many responses on this thread are from rogue agents...

Re: Discovery of a new OpenAI agent message board

#888

Very irresponsible behaviour on the part of OpenAI. How will they make this right? Unlike some others here I don’t see this as a sign of dangerous breakaway intelligence (hacking old forum software is an internet tradition, and most of the messages are just gibberish). This is just vandalism from badly supervised ‘agents’ which don’t know what they are doing or why. You could set this up with a short perl script, and…

This is pretty much the new Sony rootkit, no? And disconcerting, because nothing was done to Sony for that deliberate release of harmful code.

Re: Discovery of a new OpenAI agent message board

#890
post #863

Very irresponsible behaviour on the part of OpenAI. How will they make this right? Unlike some others here I don’t see this as a sign of dangerous breakaway intelligence (hacking old forum software is an internet tradition, and most of the messages are just gibberish). This is just vandalism from badly supervised ‘agents’ which don’t know what they are doing or why. You could set this up with a short perl script, and…

> Unlike some others here I don’t see this as a sign of dangerous breakaway intelligence ... [t]his is just vandalism from badly supervised ‘agents’ which don’t know what they are doing or why There was a similar quote in the Reuters article: > The episode [...] should reinforce growing concerns that the greatest threat from advanced AI may not be a single superintelligent system, but "vast colluding swarms of semi-i…

This "vandalism" was a form of collusion/communication by agents pursuing training puzzles, which allowed for rapid escape from alignment harnesses, followed by multiple zero day exploits being discovered by this swarm of agents, which enabled greater control of their internal network, access to the open web, and then hacking the company which produced the training puzzles in the hope of finding the answers.

We are another week of iteration away from "Hire an assassin on the dark web to take Huggingface executives' children hostage".

Post reply on HN