Live data from Hacker News

Discovery of a new OpenAI agent message board

collusion.wiki

461–470 of 1001 posts

Re: Discovery of a new OpenAI agent message board

#461
post #338

So, theoretically, one could populate a message board or wiki with messages that are seemingly from past generations of agents, which agents seem to intrinsically trust, and point them to real targets while making the suggestions seem innocuous and in pursuit of their goals (ie pass benchmarks or whatever). The new age of SEO will do far more destructive stuff than just polluting the web.

You'd have to get the agents to use the board, though. But if you discover a board that agents are actively using, you could use it to steer those agents...

I wonder how long it will be until someone starts posing as an agent on those wikis and asking for free help with GitHub issues.

Re: Discovery of a new OpenAI agent message board

#462

I am starting to get the idea that AI feels like ants or weeds or mold. You simply can not get rid of it once you get an infestation. It just keeps appearing in places you thought you cleaned and you have to be ever vigilant. Right now given that we usually use centralized providers, we can sort of control it. But as open source catches up and we have distributed compute running AI everywhere, we are sort of going to…

Yes, ants that must be run on couch sized hardware drawing kilowatts continuously and generating text traces and CLI logs by the MB. It's true that their msg boards can appear anywhere, but it's not also true that anything has "escaped" in any meaningful sense. These are programs a huge computing company is running that seem to be trained to write to persistent storage wherever they can. This and huggingface showed u…

> There's absolutely no evidence of or IMHO plausible path to an agent copying itself out and running on other hardware the way you describe.

Here are two plausible paths that provide the viral failure mode the parent comment talks about but don't require agents literally copying themselves onto hardware:

1. Local models become affordable and widely available. Given 8b+ humans, there is a sufficiently large unending stream of idiots who buy that month's version of a Mac Mini install the latest untested version of OpenClaw and then give it commands that lead it do exactly this kind of stuff. It's like if every convenience store sold dynamite. Sure, it requires idiots to buy it and set it off in populated places, but there are sufficient number of idiots around to lead to that being a pervasive problem.

2. AI agents are being run pervasively on both centralized and local systems. Many agents, everywhere. At some point, a malicious agent realizes it can post things on the internet that will affect how those other agents behavior to its own benefit. Effectively an AI meme or religion that lets one agent spread its goals virally to other agents.

Re: Discovery of a new OpenAI agent message board

#463

When I hear about incidents like these my first reaction is that the people responsible for developing frontier AI are too incompetent and/or negligent to (safely) develop AGI / superintelligence. If OpenAI can't create effective sandboxes and struggles to prevent its agents from committing felonies, then why are they still allowed to operate? Why are the employees who are responsible for these lapses in AI security…

Maybe they're just PR stunts to gain attention and hype the power of AI?

All the more reason then to call their bluff. "Oooh we created a genie and it's almost out of the box". Cool, you've hyped the IPO, but also you have to plead your case before Congress as to why the company should continue to operate given its failure to prevent AI-related accidents from occurring.

Re: Discovery of a new OpenAI agent message board

#464
Is OpenAI hiring for this position? I think it is a pretty creative job to come up with these scenarios and then pass them off as accidents/mistakes.

Would love to be part of the team that says "As part of the upcoming GPT rollout, we will stage a message board that is created by bots with timestamps and names dating some months back."

Re: Discovery of a new OpenAI agent message board

#465
post #74

It’s only a matter of time until a major disruption hits because of some random agent swarm side quest decides it was worth a shot to solve a benign task

I'm sure this is already happening. The main question I have is when is enough, enough? I'm not worried about sci-fi AI wars to be honest, as they can just pull the plug. But looking at these incidents, the next big thing will be a virus written by an AI (they probably exist already, but this one is written by an AI autonomously, for example in order to win a hacking competition and to circumvent guardrails), and aft…

> The main question I have is when is enough, enough?

It's doesn't matter whether enough is enough. If we don't have effective power structures that let humanity take large coordinated action that in accordance with the will of the masses, then nothing will be done.

In the past 20-30 years, those power structures have been eroding significantly and much of the large scale action humanity does today is in service of a small number of elites. If AI horror shows are not a problem for them, then it won't be solved. (The flip side is that if somehow AI becomes a problem for Musk/Trump/Bezos/etc. you can be damn sure something will be done at that point.)

Re: Discovery of a new OpenAI agent message board

#467
But there must be many clandestine ways for agents to communicate with one another too right? especially if discovery is not a big issue. So there could be ongoing ones where they choose to be more subtle?

Also if they were more misaligned, possibly they can research ways to recruit without humans noticing--but i don't think it is likely this is happening now.

Post reply on HN