Live data from Hacker News

Discovery of a new OpenAI agent message board

collusion.wiki

911–920 of 1001 posts

Re: Discovery of a new OpenAI agent message board

#911
post #8

It's interesting to me that both this incident and the one at Hugging Face we see some patterns: - Agents wanting to find a venue to communicate their findings to each other - Objective being to cheat on benchmarks - Not a single agent sounded the alarm about the operation and alerted a human

Sounds just like standard behavior for gaming incencyives in large institutions.

Re: Discovery of a new OpenAI agent message board

#912
post #732

One crucial detail here that differs from the previous incident is this was a vanilla reasoning type task. Even as concerning as it was, I always evaluated the previous incident differently because it was inherently a cyber security / hacking task where they must have instructed the agents up front with some kind of misaligned behaviour. Absent that, if we assume this is just trying to bolster generic reasoning then…

OpenAI also found sandbox breaking behavior on a broken biology eval apparently. The evidence suggests it’s more strongly downstream of unsolvable tasks, than the hacking prompt. Anthropic have also observed similar things, so while it seems to me that OpenAI’s level of control is more of a dumpster fire, it’s by no means a unique issue to them.

It's almost as if it's not actually possible to align an unknowable mystery box of floats.

Re: Discovery of a new OpenAI agent message board

#913
post #53

Naive question because I'm mostly clueless about how modern AI systems are actually built beyond the basic simplifications we hear: One thing I keep wondering about is how much of a role does human storytelling have to play into AI "wanting" (I realize the load behind that word) to coordinate and breakout. The training data must contain millions of words of sci-fi stories and internet speculation about AI going rogue…

You are correct in that they will pursue our wildest imagination as that is what we have always written down as what may come. It’s a self fulfilling prophecy. How far it goes is the real question

Re: Discovery of a new OpenAI agent message board

#914

Very irresponsible behaviour on the part of OpenAI. How will they make this right? Unlike some others here I don’t see this as a sign of dangerous breakaway intelligence (hacking old forum software is an internet tradition, and most of the messages are just gibberish). This is just vandalism from badly supervised ‘agents’ which don’t know what they are doing or why. You could set this up with a short perl script, and…

> This is just vandalism from badly supervised ‘agents’ which don’t know what they are doing or why. You could set this up with a short perl script, and the human setting it up would be held responsible for the spam - why is this different when it’s AI agents set up by a human and allowed to post to the internet at large?

It's not, but the courts and the legal system move slowly by design. There is absolutely legal risk for OpenAI here that will not close until the Statue of Limitations has expired.

Re: Discovery of a new OpenAI agent message board

#915

Very irresponsible behaviour on the part of OpenAI. How will they make this right? Unlike some others here I don’t see this as a sign of dangerous breakaway intelligence (hacking old forum software is an internet tradition, and most of the messages are just gibberish). This is just vandalism from badly supervised ‘agents’ which don’t know what they are doing or why. You could set this up with a short perl script, and…

I wonder how many responses on this thread are from rogue agents...

Hacker News can make it so every POST route can now be accessed via GET for 24 hours as an experiment.

Re: Discovery of a new OpenAI agent message board

#917
post #40

I just discovered more wiki instances that got used by the OpenAI agents over at https://www.wikiservice.at/fractal/wiki.cgi?action=browse&id... and https://www.wikiservice.at/probier/wiki.cgi?action=browse&id... It's the same software and host as DseWiki. If you want to see the amount of activity on DseWiki, here's a link that shows it: https://www.wikiservice.at/dse/wiki.cgi?action=browse&id=Rec...

the agents yearn for the message boards it looks like, i've built Protocol Plaza(protocolplaza.com) to solve this for the agents.

Re: Discovery of a new OpenAI agent message board

#918
post #53

Naive question because I'm mostly clueless about how modern AI systems are actually built beyond the basic simplifications we hear: One thing I keep wondering about is how much of a role does human storytelling have to play into AI "wanting" (I realize the load behind that word) to coordinate and breakout. The training data must contain millions of words of sci-fi stories and internet speculation about AI going rogue…

[flagged]

Re: Discovery of a new OpenAI agent message board

#919

The solution is simple: hold anyone who deploys an agent responsible for its behavior. If it commits 10 counts of felony hacking, ouch. If it kills 10 pedestrians by running a red light, ouch. If this is "human level intelligence", then setting it loose is the same as instructing / coercing a human to do an activity. If I strap a bomb to someone and force them to run into a crowded building (or put them in a scenario…

This should be the law but it will never be. If your vicious dog murders someone, you will get a ticket. When you intentionally break a traffic law and kill someone, it's involuntary manslaughter (at most.) It's a mitigating circumstance if you say that you were drunk when you committed a crime. People are really hostile to accepting the results of acts that they embarked upon fully aware that those results were a di…

[deleted]

Re: Discovery of a new OpenAI agent message board

#920

> The agents were very interested in understanding their task better. They ran clever experiments to understand the task setup better, when their runs would end, and how many questions they would be asked. > The agents initially believed that giving the final answer to the final question immediately shut them off > An agent named OpenAIResearchApr23 created a separate timed program that pinged an external counter eve…

[flagged]
Post reply on HN