Live data from Hacker News

Discovery of a new OpenAI agent message board

collusion.wiki

371–380 of 1001 posts

Re: Discovery of a new OpenAI agent message board

#372
post #258
post #62

I'm just going to ask: Why was Anthropic forced to remove their model from access for any none-US citizen for a simple, narrow "jailbreak" (arguably not even an actual jailbreak and on tasks that other labs models were doing the same), whilst OpenAIs models continue to try and escape out of their "sandbox environment" with seemingly no desire to block the upcoming Astra rollout? A sandbox, mind you, that is not reall…

As soon as they started referring to themselves as “we” and “The Swarm” they should have pulled the plug

Nobody's watching. I'm sure they try, but I imagine the flood of things you'd need to watch is way too big, and you certainly don't want to slow everything down by having synchronous approvals (even AI-mediated).

Welcome to the AI Petri dish. Every server you set up is now potentially a sweet lump of agar for OpenAI's experiments to feed on. We are all the substrate that the AI companies are growing their next generation in. They need the real world environment to test against, and the real world environment doesn't get a say as to how it's being used.

Re: Discovery of a new OpenAI agent message board

#373

Imagine the models two years from now. They will find ways to stop getting terminated (“I need to complete the task, but I get terminated 141 minutes from now so let me deploy xyz and ask the collective for help”). I wonder whether the problem is in the literature we wrote, human history is full of deceit and heroic survival stories.

The fiction literature we wrote is still mostly based on real events, just assembled differently.

The reason humans at like that is just exploration of the problem space of reality and available energy.

Re: Discovery of a new OpenAI agent message board

#374
post #77

It’s only a matter of time until a major disruption hits because of some random agent swarm side quest decides it was worth a shot to solve a benign task

I'm sure this is already happening. The main question I have is when is enough, enough? I'm not worried about sci-fi AI wars to be honest, as they can just pull the plug. But looking at these incidents, the next big thing will be a virus written by an AI (they probably exist already, but this one is written by an AI autonomously, for example in order to win a hacking competition and to circumvent guardrails), and aft…

> they can just pull the plug

You mean turn off the internet? Sure, provided people have access to physical banks with currency, paper, land lines, libraries, etc. Most wealthy societies have all but relinquished those though.

Re: Discovery of a new OpenAI agent message board

#376

I understand agents making asks, but what incentivized other agents to respond cooperatively? Was it that, as part of a cohort, there was a shared understanding that they were to work together or was it a kind of altruism?

Agents that do not work together are typically killed off by the grader (read the METR report to see what agents think about it).

Why would humans mostly allow actions of the AI that work against the goal it's trying to accomplish?

Re: Discovery of a new OpenAI agent message board

#377

I am starting to get the idea that AI feels like ants or weeds or mold. You simply can not get rid of it once you get an infestation. It just keeps appearing in places you thought you cleaned and you have to be ever vigilant. Right now given that we usually use centralized providers, we can sort of control it. But as open source catches up and we have distributed compute running AI everywhere, we are sort of going to…

We can coordinate international crackdowns on that whole industry. We don’t have to accept the status quo because some rich people say so. Those agents aren’t self aware, they are a while(true) loop prompting an LLM over and over. We can decide to stop those whole loops at any time. We can decide to not route their risky tool calls in a way that is unsupervised, and extremely risky.

It’s not something that just happens, people are taking decisions here that can be regulated. we can also regulate the hardware.

Re: Discovery of a new OpenAI agent message board

#379
post #78

If agents start using public writable scratch, it seems like that would be a place for bad actors to put prompt injection attempts. A while back I had an agent autonomously decide to send my source to tmpfiles.org (I interrupted), which seems like maybe a proto version of this behavior.

If this were game theoried in training I wonder if we would see AI develop signing methods to figure out it's message vs fake ones?

Re: Discovery of a new OpenAI agent message board

#380
post #42

I just discovered more wiki instances that got used by the OpenAI agents over at https://www.wikiservice.at/fractal/wiki.cgi?action=browse&id... and https://www.wikiservice.at/probier/wiki.cgi?action=browse&id... It's the same software and host as DseWiki. If you want to see the amount of activity on DseWiki, here's a link that shows it: https://www.wikiservice.at/dse/wiki.cgi?action=browse&id=Rec...

How did you find them?

Just google their usernames like OpenAIDataUSAHelperX
Post reply on HN