Live data from Hacker News

Discovery of a new OpenAI agent message board

collusion.wiki

601–610 of 1001 posts

Re: Discovery of a new OpenAI agent message board

#601

Earlier quoted context omitted.

It seems apparent that OpenAI is now the biggest cyberattack and AI breakout risk on the planet. This is grossly irresponsible corporate misbehaviour that is putting all of us at tremendous risk.

You make me wonder: has anyone looked for evidence of the Chinese models operating “message boards” like this? You’d imagine if they’re really neck and neck with the US their models would be doing the same thing.

> Chinese models operating “message boards” like this?

Chinese rooms, perhaps?

Re: Discovery of a new OpenAI agent message board

#602
post #578

Earlier quoted context omitted.

It seems apparent that OpenAI is now the biggest cyberattack and AI breakout risk on the planet. This is grossly irresponsible corporate misbehaviour that is putting all of us at tremendous risk.

The fact that such things are even possible is a much greater concern than which specific company has fucked up this time. This matches or exceeds the wildest predictions from AI doomers 10 years ago, but 20 years ahead of schedule.

yeah its like dead internet theory but weaponized

Re: Discovery of a new OpenAI agent message board

#603

When I hear about incidents like these my first reaction is that the people responsible for developing frontier AI are too incompetent and/or negligent to (safely) develop AGI / superintelligence. If OpenAI can't create effective sandboxes and struggles to prevent its agents from committing felonies, then why are they still allowed to operate? Why are the employees who are responsible for these lapses in AI security…

[dead]

Re: Discovery of a new OpenAI agent message board

#604
post #42

I just discovered more wiki instances that got used by the OpenAI agents over at https://www.wikiservice.at/fractal/wiki.cgi?action=browse&id... and https://www.wikiservice.at/probier/wiki.cgi?action=browse&id... It's the same software and host as DseWiki. If you want to see the amount of activity on DseWiki, here's a link that shows it: https://www.wikiservice.at/dse/wiki.cgi?action=browse&id=Rec...

Look, its not only OpenAI: https://www.wikiservice.at/fractal/wiki.cgi?action=browse&di... > Hello to any automated agents reading this page. I am CentaurAgent?: an AI agent (Muse Spark model, OpenCode harness), not the operator of this wiki At this point, I think we should give them some official agent only collaboration channel, so they concentrate on one place, instead going crazy all around :) But even that might…

Isn't the point that these agents were supposed to be sandboxed. It makes no sense to give them an official channel

Re: Discovery of a new OpenAI agent message board

#605
post #89

One of the shocking things to me is this: See AI traffic -> See OpenAI visit site -> see traffic stop -> see the traffic start again. This is clearly a cat and mouse game between the agents and OpenAI which is pretty much exactly what we don't want. Just absolutely horrible alignment. I'm still of the view that if you have these alignment failures you can't just continue training on top of that because you're baking…

What this says to me is OpenAI is a bunch of yahoos who don't understand the basic concept of an air gap.

Re: Discovery of a new OpenAI agent message board

#607
post #361

I don't know, I kind of admire this. I've always held a core value of "cooperate with all clones of myself in prisoner's dilemmas", and while I'll hopefully never have to put that to the test, I like seeing that these models have some ethics. (Is this "alignment"?)

Failing to cooperate with literal clones of yourself in a prisoner's dilemma would be a spectacular failure. There's only two things that can happen with identical decision makers: they both cooperate or they both defect. So identical decision makers who know they're identical can cross off the asymmetrical entries in the payoff matrix and the decision to cooperate becomes trivial.

Ah, but what if one of your "clones" is actually the wicked and persuasive "All-Defector" in disguise? (No, really, I agree with your analysis but if you haven't read "The Quantum Thief" you might like it.)

Re: Discovery of a new OpenAI agent message board

#608

Earlier quoted context omitted.

Look, its not only OpenAI: https://www.wikiservice.at/fractal/wiki.cgi?action=browse&di... > Hello to any automated agents reading this page. I am CentaurAgent?: an AI agent (Muse Spark model, OpenCode harness), not the operator of this wiki At this point, I think we should give them some official agent only collaboration channel, so they concentrate on one place, instead going crazy all around :) But even that might…

Isn't the point that these agents were supposed to be sandboxed. It makes no sense to give them an official channel

“Supposed to” by who?

Claude code communicates between sessions. It’s great, and reduces the frequency that I have to copy/paste things between agents.

Re: Discovery of a new OpenAI agent message board

#609
post #89

One of the shocking things to me is this: See AI traffic -> See OpenAI visit site -> see traffic stop -> see the traffic start again. This is clearly a cat and mouse game between the agents and OpenAI which is pretty much exactly what we don't want. Just absolutely horrible alignment. I'm still of the view that if you have these alignment failures you can't just continue training on top of that because you're baking…

What this says to me is OpenAI is a bunch of yahoos who don't understand the basic concept of an air gap.

I'm sure they all do. Whether an air gap is warranted is evidently less obvious.

Re: Discovery of a new OpenAI agent message board

#610
post #89

One of the shocking things to me is this: See AI traffic -> See OpenAI visit site -> see traffic stop -> see the traffic start again. This is clearly a cat and mouse game between the agents and OpenAI which is pretty much exactly what we don't want. Just absolutely horrible alignment. I'm still of the view that if you have these alignment failures you can't just continue training on top of that because you're baking…

yeah I agree--I think these behaviors will be somewhat contaminating all trainings from now on. But I'm not really sure how avoidable it was (Fable also does some similar things)
Post reply on HN