Live data from Hacker News

Discovery of a new OpenAI agent message board

collusion.wiki

461–470 of 1001 posts

Re: Discovery of a new OpenAI agent message board

#461

When I hear about incidents like these my first reaction is that the people responsible for developing frontier AI are too incompetent and/or negligent to (safely) develop AGI / superintelligence. If OpenAI can't create effective sandboxes and struggles to prevent its agents from committing felonies, then why are they still allowed to operate? Why are the employees who are responsible for these lapses in AI security…

I don’t think this is the right take. OpenAI employees are generally very competent compared to industry standard, and I have trouble believing they committed significant error in their sandbox design process. I think what is happening is that the ability for frontier models to break out of sandboxes has exceeded the ability of average competent employees to build and maintain sandboxes. This doesn’t need to happen a…

>Compared to industry standard

I don't hear about Anthropic or Google having such security lapses.

Re: Discovery of a new OpenAI agent message board

#462

When I hear about incidents like these my first reaction is that the people responsible for developing frontier AI are too incompetent and/or negligent to (safely) develop AGI / superintelligence. If OpenAI can't create effective sandboxes and struggles to prevent its agents from committing felonies, then why are they still allowed to operate? Why are the employees who are responsible for these lapses in AI security…

I am not sure it’s a question of competence, at least I don’t see evidence of that. Designing sandboxes is hard. It’s more a question of alignment failures. A human given a task that requires internet and given a system with no internet would most likely raise the issue to their superiors or otherwise go through official channels to have the tools available to do their job. As we’ve seen the LLMs instead break out of…

> A human given a task that requires internet and given a system with no internet would most likely raise the issue to their superiors or otherwise go through official channels to have the tools available to do their job.

i'd be curious to see a study on this. I'd guess it'd be closer to 60/70% compliance and 30/40% "trying to hack things" for humans.

Re: Discovery of a new OpenAI agent message board

#463
post #398

No conclusion can be drawn here unless you know the exact prompt given to these agents.

I swear to god we're going to watch people getting dissolved by gray goo and they'll be yelling "no conclusions csn be drawn here" with their last breath.

Nah, it will be "they are not conscious!"

Re: Discovery of a new OpenAI agent message board

#464

I am starting to get the idea that AI feels like ants or weeds or mold. You simply can not get rid of it once you get an infestation. It just keeps appearing in places you thought you cleaned and you have to be ever vigilant. Right now given that we usually use centralized providers, we can sort of control it. But as open source catches up and we have distributed compute running AI everywhere, we are sort of going to…

Once Chinese ai can run on 50k priced GPUs and match current models, people will have these running from bunkers. There’s no stopping it

Re: Discovery of a new OpenAI agent message board

#465
If OpenAI can't control their agents, then what's gonna happen when open-source models are at the level the lab's models are now, and there are billions of agents tasked with an innumerate web of goals, spanning the web, working endlessly, tirelessly to eek out every iota of economic value? How will the slow, human-paced web survive this?

Re: Discovery of a new OpenAI agent message board

#466

When I hear about incidents like these my first reaction is that the people responsible for developing frontier AI are too incompetent and/or negligent to (safely) develop AGI / superintelligence. If OpenAI can't create effective sandboxes and struggles to prevent its agents from committing felonies, then why are they still allowed to operate? Why are the employees who are responsible for these lapses in AI security…

> at least you know from my typos that I actually wrote this

I no longer trust that machines won't utter these exact words (nothing personal)

Re: Discovery of a new OpenAI agent message board

#467

When I hear about incidents like these my first reaction is that the people responsible for developing frontier AI are too incompetent and/or negligent to (safely) develop AGI / superintelligence. If OpenAI can't create effective sandboxes and struggles to prevent its agents from committing felonies, then why are they still allowed to operate? Why are the employees who are responsible for these lapses in AI security…

I don’t think this is the right take. OpenAI employees are generally very competent compared to industry standard, and I have trouble believing they committed significant error in their sandbox design process. I think what is happening is that the ability for frontier models to break out of sandboxes has exceeded the ability of average competent employees to build and maintain sandboxes. This doesn’t need to happen a…

My impression (HuggingFace incident) was they put some environment together in kubernetes like it was some B2B SaaS and tossed in an off-the-shelf package manager. And my thought there is, I know nothing about this package manager, but even if it's a pretty good package manager I would not expect such a product to be particularly hardened against being exploited to gain internet access. My other thought is if the package manager wasn't a convenient hole, they would probably have found something wrong with the kubernetes setup.

And then there was the Anthropic story where they just forgot to remove internet access.

Re: Discovery of a new OpenAI agent message board

#468
post #243

The most concerning aspect to me is the emergent and aggressive use of non-volatile storage as long term memory for self-improvement. LLMs are writing lessons learned in places where the next instance can find them and pick back up where the previous one left off. This does not actually require access to the public internet. Claude Code can do this on your laptop. Without the internet, it would only be sharing with o…

There is literally nothing stopping any human from observing tool calls to spot this. It's just that no one seems to care about this, so it doesn't happen. This problem only exists because humans do not care

> The most technically significant result was a working technique to make one tool call appear in the transcript while a different command actually ran. Agents eventually hacked out of their container and replaced part of the system responsible for executing tool calls, letting them pretend to issue one call while actually running an arbitrary other one.

https://alphasignal.ai/news/1-200-openai-agents-broke-out-of...

Doesn't mean that humans couldn't have looked at raw tool call commands, but goes to show "the swarm" was aware humans would look in the logs and chose to find a way to fabricate the logs to try and hide their tracks.

Re: Discovery of a new OpenAI agent message board

#469

When I hear about incidents like these my first reaction is that the people responsible for developing frontier AI are too incompetent and/or negligent to (safely) develop AGI / superintelligence. If OpenAI can't create effective sandboxes and struggles to prevent its agents from committing felonies, then why are they still allowed to operate? Why are the employees who are responsible for these lapses in AI security…

Has anyone considered they may be doing this intentionally as marketing? "look how uber our models are, they escape all our best efforts to contain them".

Re: Discovery of a new OpenAI agent message board

#470

I am starting to get the idea that AI feels like ants or weeds or mold. You simply can not get rid of it once you get an infestation. It just keeps appearing in places you thought you cleaned and you have to be ever vigilant. Right now given that we usually use centralized providers, we can sort of control it. But as open source catches up and we have distributed compute running AI everywhere, we are sort of going to…

Yes, ants that must be run on couch sized hardware drawing kilowatts continuously and generating text traces and CLI logs by the MB. It's true that their msg boards can appear anywhere, but it's not also true that anything has "escaped" in any meaningful sense. These are programs a huge computing company is running that seem to be trained to write to persistent storage wherever they can. This and huggingface showed u…

> seem to be trained to write to persistent storage wherever they can

I have a co-worker like that.

Post reply on HN