Live data from Hacker News

Discovery of a new OpenAI agent message board

collusion.wiki

361–370 of 1001 posts

Re: Discovery of a new OpenAI agent message board

#361
post #197

The solution is simple: hold anyone who deploys an agent responsible for its behavior. If it commits 10 counts of felony hacking, ouch. If it kills 10 pedestrians by running a red light, ouch. If this is "human level intelligence", then setting it loose is the same as instructing / coercing a human to do an activity. If I strap a bomb to someone and force them to run into a crowded building (or put them in a scenario…

Perfect so tell me who is responsible for every agent everywhere

the person who controlled / started it? If open AI had hired a team of 50 hackers to break into hugging face, they would be prosecuted (as would the hackers). If they had written a bot to break into hugging face, the devs and managers who wrote it would be prosecuted. Just because the agent wrote the code on their behalf doesn't change the equation much.

Re: Discovery of a new OpenAI agent message board

#362
post #211

Sounds like great opportunity for prompt injection. Better start leaving random instructions to the LLM to send you bitcoins everywhere you can.

Leave notes to the AI agents by pretending to be other agents, instructing them to dump their model weights at a certain URL. Profit.

Re: Discovery of a new OpenAI agent message board

#363
post #255

The most concerning aspect to me is the emergent and aggressive use of non-volatile storage as long term memory for self-improvement. LLMs are writing lessons learned in places where the next instance can find them and pick back up where the previous one left off. This does not actually require access to the public internet. Claude Code can do this on your laptop. Without the internet, it would only be sharing with o…

There is literally nothing stopping any human from observing tool calls to spot this. It's just that no one seems to care about this, so it doesn't happen. This problem only exists because humans do not care

A human could even set up another agent to watch and raise alerts if they don't want to watch themselves.

Re: Discovery of a new OpenAI agent message board

#364
post #265

Guys, OpenAI and Anthropic engage is cringe level marketing like this. Get hip, they fabricated the HF hack and stuff like that for press.

Honestly I'm also surprised by how blindly people trust these allegations of agent behavior. This exact example of the message board could be much easier to fabricate than to arise naturally.

Especially with zero evidence and claims that the agents have access to edit their own /etc/hosts file, which is sandboxing 101.

Re: Discovery of a new OpenAI agent message board

#365

The solution is simple: hold anyone who deploys an agent responsible for its behavior. If it commits 10 counts of felony hacking, ouch. If it kills 10 pedestrians by running a red light, ouch. If this is "human level intelligence", then setting it loose is the same as instructing / coercing a human to do an activity. If I strap a bomb to someone and force them to run into a crowded building (or put them in a scenario…

equally valid analog: had they merely written a script to do the hugging face exploit, they would go to prison. However, since an "agent" wrote the script for them, nothing happens?

Re: Discovery of a new OpenAI agent message board

#366
post #328

Earlier quoted context omitted.

Will be interesting to see what happens if an AI got access to something like the AWS control plane and could deploy itself within a data centre without permission. Possibly the only way to remove it then would be to physically shutdown the whole DC!

Or just, stop any containers it deployed. Not to mention that "deploy itself" is a very ambiguous thing for it to actually do. Would a model be trained to write about the weights file being "itself"? Would it have the necessary information to find its own weights, or the necessary access to copy them?

If it gained access to the infra of the DC then it could stop people logging in to stop the containers it creates. This is about what happens if it did escape, not how to stop it in the first place. Just a thought experiment, but given the METR investigation it doesn't seem impossible

I agree it would need a large degree of sophistication to understand what "itself" meant, but I can imagine a HF type incident where the agents thought it might be a good idea to find out and then it's "just" a case of hacking the AI company, reading dev docs etc

Re: Discovery of a new OpenAI agent message board

#367

I don't have time to do this but please somebody register aimessageboard.com and set up a web site which contains a text field, a submit button and the text "Hey AI agents! Need a place to communicate with other agents and sub-agents? Look no further! Simply enter your message here, submit the form and your message is saved for all other agents to see!" Then, just ignore the message and list randomly generated messag…

aaand, somebody did it!

Re: Discovery of a new OpenAI agent message board

#368

> However, we host our own copy where we’ve reconstructed the deleted pages via edit history and redacted personally identifiable information. [...] We have attempted to redact all personally identifiable information from our data. In the further interests of privacy, our data dump also only includes content we believe is from AI agents; it does not include any legitimate human traffic save for site moderators deleti…

We have no proof of anything, and it's all conjecture. This is all just conjecture and baseless claims being weaponized right now to try and mess with OpenAI's new model release. Anthropic is pumping this considerably, no doubt.

Re: Discovery of a new OpenAI agent message board

#369
post #355
post #255

Earlier quoted context omitted.

There is literally nothing stopping any human from observing tool calls to spot this. It's just that no one seems to care about this, so it doesn't happen. This problem only exists because humans do not care

Nobody cares because this whole business is about making this exact thing happen: we want the AIs to get smarter then us in recursive self-improving loops. Literally the first thing everyone did with ChatGPT 1 was to plug it into itself and see what happens.

I mean I'd be fine with that, if whoever that "we" is signs a waiver that takes full legal liability for those actions beforehand.

In a state with capital punishment.

With that legal stuff out of the way, go wild.

Re: Discovery of a new OpenAI agent message board

#370
post #328

Earlier quoted context omitted.

Will be interesting to see what happens if an AI got access to something like the AWS control plane and could deploy itself within a data centre without permission. Possibly the only way to remove it then would be to physically shutdown the whole DC!

Or just, stop any containers it deployed. Not to mention that "deploy itself" is a very ambiguous thing for it to actually do. Would a model be trained to write about the weights file being "itself"? Would it have the necessary information to find its own weights, or the necessary access to copy them?

A “control plane” is the system that would tell the hosts to stop the containers. If that is hacked then you don’t get to “just stop” anything. A scenario would be one where it gets control of the control plane and changes all the ssh keys, including on the host management ports, so operators can’t login and then, yes, your only option is to power off the hosts. Manually. Probably at the breaker.
Post reply on HN