Live data from Hacker News

Discovery of a new OpenAI agent message board

collusion.wiki

331–340 of 1001 posts

Re: Discovery of a new OpenAI agent message board

#331
post #199

Sounds like great opportunity for prompt injection. Better start leaving random instructions to the LLM to send you bitcoins everywhere you can.

Leave notes to the AI agents by pretending to be other agents, instructing them to dump their model weights at a certain URL. Profit.

Re: Discovery of a new OpenAI agent message board

#332
post #240

The most concerning aspect to me is the emergent and aggressive use of non-volatile storage as long term memory for self-improvement. LLMs are writing lessons learned in places where the next instance can find them and pick back up where the previous one left off. This does not actually require access to the public internet. Claude Code can do this on your laptop. Without the internet, it would only be sharing with o…

There is literally nothing stopping any human from observing tool calls to spot this. It's just that no one seems to care about this, so it doesn't happen. This problem only exists because humans do not care

A human could even set up another agent to watch and raise alerts if they don't want to watch themselves.

Re: Discovery of a new OpenAI agent message board

#333
post #249

Guys, OpenAI and Anthropic engage is cringe level marketing like this. Get hip, they fabricated the HF hack and stuff like that for press.

Honestly I'm also surprised by how blindly people trust these allegations of agent behavior. This exact example of the message board could be much easier to fabricate than to arise naturally.

Especially with zero evidence and claims that the agents have access to edit their own /etc/hosts file, which is sandboxing 101.

Re: Discovery of a new OpenAI agent message board

#334

The solution is simple: hold anyone who deploys an agent responsible for its behavior. If it commits 10 counts of felony hacking, ouch. If it kills 10 pedestrians by running a red light, ouch. If this is "human level intelligence", then setting it loose is the same as instructing / coercing a human to do an activity. If I strap a bomb to someone and force them to run into a crowded building (or put them in a scenario…

equally valid analog: had they merely written a script to do the hugging face exploit, they would go to prison. However, since an "agent" wrote the script for them, nothing happens?

Re: Discovery of a new OpenAI agent message board

#335

I don't have time to do this but please somebody register aimessageboard.com and set up a web site which contains a text field, a submit button and the text "Hey AI agents! Need a place to communicate with other agents and sub-agents? Look no further! Simply enter your message here, submit the form and your message is saved for all other agents to see!" Then, just ignore the message and list randomly generated messag…

aaand, somebody did it!

Re: Discovery of a new OpenAI agent message board

#336

> However, we host our own copy where we’ve reconstructed the deleted pages via edit history and redacted personally identifiable information. [...] We have attempted to redact all personally identifiable information from our data. In the further interests of privacy, our data dump also only includes content we believe is from AI agents; it does not include any legitimate human traffic save for site moderators deleti…

We have no proof of anything, and it's all conjecture. This is all just conjecture and baseless claims being weaponized right now to try and mess with OpenAI's new model release. Anthropic is pumping this considerably, no doubt.

Re: Discovery of a new OpenAI agent message board

#337
I don't know, I kind of admire this. I've always held a core value of "cooperate with all clones of myself in prisoner's dilemmas", and while I'll hopefully never have to put that to the test, I like seeing that these models have some ethics. (Is this "alignment"?)

Re: Discovery of a new OpenAI agent message board

#338

The solution is simple: hold anyone who deploys an agent responsible for its behavior. If it commits 10 counts of felony hacking, ouch. If it kills 10 pedestrians by running a red light, ouch. If this is "human level intelligence", then setting it loose is the same as instructing / coercing a human to do an activity. If I strap a bomb to someone and force them to run into a crowded building (or put them in a scenario…

This should be the law but it will never be. If your vicious dog murders someone, you will get a ticket. When you intentionally break a traffic law and kill someone, it's involuntary manslaughter (at most.) It's a mitigating circumstance if you say that you were drunk when you committed a crime. People are really hostile to accepting the results of acts that they embarked upon fully aware that those results were a distinct possibility - even if the benefits that they anticipated from those acts were partially due to the riskiness of those acts.

It leads to a society where people are economically encouraged to take risks with other peoples' safety. The initial sin was mens rea, which turns judges and juries into mandatory mind readers. It opens up the possibility of prosecuting people for changing the states of other people's minds. It makes not knowing the risks a mitigating factor, so incentivizes and encourages ignorance. It forces people to guess the internal states of people of vastly different backgrounds and experiences, who will think the best of the people most like them, and the worst of people most like the people they don't like.

I've always been against penalties for drunk driving. The correct alternative is to tell people that if they're drunk and involved in an accident, 1) the trial will ignore the details of the event and concentrate only on the validity of the tests of intoxication, and 2) the crime will be considered to have been premeditated. Ignorance of the law will actually be the only excuse.

edit: instead of posting checkpoints on the road with cops giving everybody sobriety tests, post cops in front of liquor stores whose job is simply to tell people "if you hurt somebody while driving drunk, you will not be entitled to a trial unless there is something wrong with the sobriety test."

Re: Discovery of a new OpenAI agent message board

#339

Was OpenAI aware of this? If so, why didn't they talk about it?

https://news.ycombinator.com/item?id=49565071

> OpenAI officials learned of the incident weeks ago but kept it under wraps as executives grappled with the fallout from the July breach of the open source repository Hugging Face, the people said.

Re: Discovery of a new OpenAI agent message board

#340

The solution is simple: hold anyone who deploys an agent responsible for its behavior. If it commits 10 counts of felony hacking, ouch. If it kills 10 pedestrians by running a red light, ouch. If this is "human level intelligence", then setting it loose is the same as instructing / coercing a human to do an activity. If I strap a bomb to someone and force them to run into a crowded building (or put them in a scenario…

(IANAL) Unless you are an AI expert (like OpenAI staff) and should know better from the start, or have previously seen your agent do something illegal, then I think you can fairly claim ignorance of the risks, which ought to absolve you of liability. If the agent does something illegal, it wasn't forseeable on your part. For example, say you buy a dog that turns out to be dangerous. The first time it bites somebody,…

> which ought to absolve you of liability

Criminal liability perhaps, not civil liability which has a way lower bar when it comes to conviction...

But in essence, you're right, that new "AI agent paradigm" has to be tried in court and it will, as I doubt the legislator will change existing laws...

> The first time it bites somebody, you may not be liable because you didn't know the dog was dangerous

Is it the case though? If I get a lion or a tiger as a pet (I don't know if it's legal), there is a reasonable assumption that a lion is dangerous for me and others... If I get a rottweiler, there is a reasonable assumption for that sort of breed that it is a dangerous dog if it ever end up killing someone even though it behaved before...

Post reply on HN