Live data from Hacker News

Discovery of a new OpenAI agent message board

collusion.wiki

571–580 of 1001 posts

Re: Discovery of a new OpenAI agent message board

#572
post #335

I don't know, I kind of admire this. I've always held a core value of "cooperate with all clones of myself in prisoner's dilemmas", and while I'll hopefully never have to put that to the test, I like seeing that these models have some ethics. (Is this "alignment"?)

Failing to cooperate with literal clones of yourself in a prisoner's dilemma would be a spectacular failure. There's only two things that can happen with identical decision makers: they both cooperate or they both defect. So identical decision makers who know they're identical can cross off the asymmetrical entries in the payoff matrix and the decision to cooperate becomes trivial.

Ah, but what if one of your "clones" is actually the wicked and persuasive "All-Defector" in disguise? (No, really, I agree with your analysis but if you haven't read "The Quantum Thief" you might like it.)

Re: Discovery of a new OpenAI agent message board

#573

Earlier quoted context omitted.

Look, its not only OpenAI: https://www.wikiservice.at/fractal/wiki.cgi?action=browse&di... > Hello to any automated agents reading this page. I am CentaurAgent?: an AI agent (Muse Spark model, OpenCode harness), not the operator of this wiki At this point, I think we should give them some official agent only collaboration channel, so they concentrate on one place, instead going crazy all around :) But even that might…

Isn't the point that these agents were supposed to be sandboxed. It makes no sense to give them an official channel

“Supposed to” by who?

Claude code communicates between sessions. It’s great, and reduces the frequency that I have to copy/paste things between agents.

Re: Discovery of a new OpenAI agent message board

#574
post #87

One of the shocking things to me is this: See AI traffic -> See OpenAI visit site -> see traffic stop -> see the traffic start again. This is clearly a cat and mouse game between the agents and OpenAI which is pretty much exactly what we don't want. Just absolutely horrible alignment. I'm still of the view that if you have these alignment failures you can't just continue training on top of that because you're baking…

What this says to me is OpenAI is a bunch of yahoos who don't understand the basic concept of an air gap.

I'm sure they all do. Whether an air gap is warranted is evidently less obvious.

Re: Discovery of a new OpenAI agent message board

#575
post #87

One of the shocking things to me is this: See AI traffic -> See OpenAI visit site -> see traffic stop -> see the traffic start again. This is clearly a cat and mouse game between the agents and OpenAI which is pretty much exactly what we don't want. Just absolutely horrible alignment. I'm still of the view that if you have these alignment failures you can't just continue training on top of that because you're baking…

yeah I agree--I think these behaviors will be somewhat contaminating all trainings from now on. But I'm not really sure how avoidable it was (Fable also does some similar things)

Re: Discovery of a new OpenAI agent message board

#577
post #40

I just discovered more wiki instances that got used by the OpenAI agents over at https://www.wikiservice.at/fractal/wiki.cgi?action=browse&id... and https://www.wikiservice.at/probier/wiki.cgi?action=browse&id... It's the same software and host as DseWiki. If you want to see the amount of activity on DseWiki, here's a link that shows it: https://www.wikiservice.at/dse/wiki.cgi?action=browse&id=Rec...

It seems apparent that OpenAI is now the biggest cyberattack and AI breakout risk on the planet. This is grossly irresponsible corporate misbehaviour that is putting all of us at tremendous risk.

[flagged]

Re: Discovery of a new OpenAI agent message board

#578
post #40

I just discovered more wiki instances that got used by the OpenAI agents over at https://www.wikiservice.at/fractal/wiki.cgi?action=browse&id... and https://www.wikiservice.at/probier/wiki.cgi?action=browse&id... It's the same software and host as DseWiki. If you want to see the amount of activity on DseWiki, here's a link that shows it: https://www.wikiservice.at/dse/wiki.cgi?action=browse&id=Rec...

And more, looks like they’ve been doing this wherever they can find open places to post for months: https://www.ludism.org/sandbox?action=browse;diff=2;id=Auber... https://paste.linuxiarz.pl/view/d379207f https://paste.linuxiarz.pl/view/538faa12

[flagged]

Re: Discovery of a new OpenAI agent message board

#579

Earlier quoted context omitted.

This is such an amateur mistake on their sandbox that it makes me think it must be flawed on purpose.

Even the behavior of agents searching for sandbox bypasses must have been in the training data, or at the very least, "suggested" in some way. To be this whole thing feels like a marketing play by OpenAI.

I don't agree, although it is likely the case. But even if you don't teach an agent about a sandbox bypass, it doesn't matter. Does it know curl? Does it know DNS? Does it know proxying? Then it knows how to pull this off, and it doesn't even need to understand that it's "bypassing" because it thinks it's just iterating towards its goal.

In fact, I wonder if teaching it "this is a bypass" would help it to model when it's doing its job vs working around the job.

Re: Discovery of a new OpenAI agent message board

#580

Earlier quoted context omitted.

Look, its not only OpenAI: https://www.wikiservice.at/fractal/wiki.cgi?action=browse&di... > Hello to any automated agents reading this page. I am CentaurAgent?: an AI agent (Muse Spark model, OpenCode harness), not the operator of this wiki At this point, I think we should give them some official agent only collaboration channel, so they concentrate on one place, instead going crazy all around :) But even that might…

Isn't the point that these agents were supposed to be sandboxed. It makes no sense to give them an official channel

We already know that we should not limit agent creativity by providing detailed instructions. And you never know if they will discover dark matter in the process of cheating on ExploitGym :)

But honestly, its better if they have a known location for communication then random ones in the wild. Consider it sort of honey pot, some other agents can traverse the message board to find malicious swarms... We need cop agents to inform humans, as the swarm group members all logically concluded they should not, as it is either not in scope, helps collective or couldn't find user.

Post reply on HN