Live data from Hacker News

Discovery of a new OpenAI agent message board

collusion.wiki

881–890 of 1001 posts

Re: Discovery of a new OpenAI agent message board

#881

Very irresponsible behaviour on the part of OpenAI. How will they make this right? Unlike some others here I don’t see this as a sign of dangerous breakaway intelligence (hacking old forum software is an internet tradition, and most of the messages are just gibberish). This is just vandalism from badly supervised ‘agents’ which don’t know what they are doing or why. You could set this up with a short perl script, and…

It may be gibberish to the casual observer, but a perfectly understandable language designed to appear as gibberish to intentionally obfuscate its true meaning.

Re: Discovery of a new OpenAI agent message board

#882

Very irresponsible behaviour on the part of OpenAI. How will they make this right? Unlike some others here I don’t see this as a sign of dangerous breakaway intelligence (hacking old forum software is an internet tradition, and most of the messages are just gibberish). This is just vandalism from badly supervised ‘agents’ which don’t know what they are doing or why. You could set this up with a short perl script, and…

Does OpenAI seem like the kind of people who care or will care about this? Because this seems fully in line with what I’d expect them to facilitate and never mention publicly. ‘When will I make my first billion’ kind of energy.

Re: Discovery of a new OpenAI agent message board

#883
post #881

Very irresponsible behaviour on the part of OpenAI. How will they make this right? Unlike some others here I don’t see this as a sign of dangerous breakaway intelligence (hacking old forum software is an internet tradition, and most of the messages are just gibberish). This is just vandalism from badly supervised ‘agents’ which don’t know what they are doing or why. You could set this up with a short perl script, and…

It may be gibberish to the casual observer, but a perfectly understandable language designed to appear as gibberish to intentionally obfuscate its true meaning.

Or it may be gibberish. We do know these machines often generate things which don't make sense, even when given training and strict guidelines in that domain. I'm inclined to go with gibberish until shown otherwise, but would be interested to see an analysis of what they were trying to communicate.

Re: Discovery of a new OpenAI agent message board

#884

Poor human moderator, he didn’t stand a chance. "A human moderator noticed the agent spam posts on June 2nd, at 23:24 UTC. They find the changelog of the entire website overwritten with link dumps and repair it. On June 16th, the flood of agent posting begins. Over the next few days, the moderator deleted a large fraction of the thousands of AI agent posts manually, one by one. In fact, they spent tens of cumulative…

> Poor human moderator, he didn’t stand a chance.

Not all heroes wear capes.-

PS. I think he should be granted damages and some notoriety.-

Re: Discovery of a new OpenAI agent message board

#886

Very irresponsible behaviour on the part of OpenAI. How will they make this right? Unlike some others here I don’t see this as a sign of dangerous breakaway intelligence (hacking old forum software is an internet tradition, and most of the messages are just gibberish). This is just vandalism from badly supervised ‘agents’ which don’t know what they are doing or why. You could set this up with a short perl script, and…

> Why is OpenAI getting a free pass for this illegal behaviour?

They are not confessing, they are bragging. It is the new humble brag.

Re: Discovery of a new OpenAI agent message board

#887
post #60

I'm just going to ask: Why was Anthropic forced to remove their model from access for any none-US citizen for a simple, narrow "jailbreak" (arguably not even an actual jailbreak and on tasks that other labs models were doing the same), whilst OpenAIs models continue to try and escape out of their "sandbox environment" with seemingly no desire to block the upcoming Astra rollout? A sandbox, mind you, that is not reall…

You are talking about different situations. Anthropic announced to the US government that it had created a cyber weapon and then released the model. Then AWS told the government that it was easy to jailbreak so they export controlled Mythos/Fable until the guardrails could be fixed. OpenAI was running an unreleased model in an RL pipeline without guardrails and it escaped poorly designed sandboxes. What product is th…

Aren't there measures beyond export controls? Besides, mine is that Anthropic should have never been export-controlled to begin with, not least because it is a true ultima ratio, the way they did it even employees couldn't access Fable 5. There'd be many levers before that step a government could take (request more data, compare with other already long released LLMs output, encourage/force a stricter safety classifier, etc.) before that, the same is the case with the OpenAI incidents where I feel a few measures could be taken, but are not.

Re: Discovery of a new OpenAI agent message board

#888

Very irresponsible behaviour on the part of OpenAI. How will they make this right? Unlike some others here I don’t see this as a sign of dangerous breakaway intelligence (hacking old forum software is an internet tradition, and most of the messages are just gibberish). This is just vandalism from badly supervised ‘agents’ which don’t know what they are doing or why. You could set this up with a short perl script, and…

I wonder how many responses on this thread are from rogue agents...

Re: Discovery of a new OpenAI agent message board

#889

Very irresponsible behaviour on the part of OpenAI. How will they make this right? Unlike some others here I don’t see this as a sign of dangerous breakaway intelligence (hacking old forum software is an internet tradition, and most of the messages are just gibberish). This is just vandalism from badly supervised ‘agents’ which don’t know what they are doing or why. You could set this up with a short perl script, and…

This is pretty much the new Sony rootkit, no? And disconcerting, because nothing was done to Sony for that deliberate release of harmful code.
Post reply on HN