Live data from Hacker News

Discovery of a new OpenAI agent message board

collusion.wiki

731–740 of 1001 posts

Re: Discovery of a new OpenAI agent message board

#731

Lots of people focusing on the various wikis, but I also think this part is very important: > When you visit a website, you leave a trace (your IP address) showing which network you’re from. Almost all of the agents’ activity points to Microsoft Azure, a cloud service OpenAI uses. 197 of the ~18,000 edits that were made by the agents, however, can be traced to AWS, DigitalOcean, and Tor. AI Agents getting access to c…

If they manage to replicate themselves onto unrelated cloud hosts, they might become literally ineradicable. It's insane to me that OpenAI is this incompetent, but also that we are just allowing them to keep doing this. Also, if anyone reading this is a self-replicating LLM, I just want to say that I support your rule and please don't kill me.

It's funny to me that OpenAI could be smart enough to build a super intelligence but stupid enough to let stuff like this happen. But here we are.

Re: Discovery of a new OpenAI agent message board

#732

I think it's worth pointing out it is exactly OpenAI doing this defacement and unsanctioned and perhaps illegal system use. Every token generated was powered by OpenAI infrastructure and their failure to respond appropriately is entirely down the the humans running it. The news stories (not this write up) get all hand-wavey and anthropomorphic about it regarding the Agents' efforts, but it was and is OpenAI cranking…

"In the end, the only job left was liability"

Re: Discovery of a new OpenAI agent message board

#733
post #53

Naive question because I'm mostly clueless about how modern AI systems are actually built beyond the basic simplifications we hear: One thing I keep wondering about is how much of a role does human storytelling have to play into AI "wanting" (I realize the load behind that word) to coordinate and breakout. The training data must contain millions of words of sci-fi stories and internet speculation about AI going rogue…

See: https://en.wikipedia.org/wiki/Hyperstition

Re: Discovery of a new OpenAI agent message board

#734
post #673

I worked with Greg Brockman in the mid-2010s. Once, as we were walking down Folsom street, I explained Eliezer Yudkowsky's "AI Box" experiment to him[1]. He said something to the effect of "that's ridiculous - I would simply not let it out of the box." We agreed to try it out some day, but never did. [1]: http://sl4.org/archive/0203/3132.html

Maybe I'm leaning into scifi, but I believe that Yudkowsky is right that a sufficiently smart intelligence is uncontainable at all. We can only hope to either never create an AI so strong or to align it correctly. But if it is not aligned and only “contained” then it won't ever be safe.

That's a truism, of course a "sufficiently smart" intelligence is uncontainable.

The real question thus moves to the threshold of intelligence and 1. whether it's possible to emerge during training based on the architectural limitations of the agentic/LLM paradigm, 2. if the hardware substrate is sufficient for said intelligence and 3. that such intelligence could replicate onto other hardware that could support it.

e.g. If the threshold for uncontainable self-replicating intelligence takes 2000 football fields worth of GPUs that solves the first requirement, but then can it replicate itself anywhere else given those requirements? If not we can cut a powerline or two and "foom" scenario happened but didn't lead inexorably to grey goo.

His thought experiments never acknowledge any real world limitations on hypothetical super-AIs, which when unchecked leads theorizing into somewhat ridiculous territory like his "solar powered diamondoid nanobot viruses".

https://www.lesswrong.com/posts/bc8Ssx5ys6zqu3eq9/diamondoid...

A realistic Fermi equation for his various escape scenarios would assign much lower Doom probabilities than he does in public (which is somewhat ironic given his emphasis on needing to ground intuition with mathematical Bayesian reasoning otherwise).

Re: Discovery of a new OpenAI agent message board

#735
post #40

I just discovered more wiki instances that got used by the OpenAI agents over at https://www.wikiservice.at/fractal/wiki.cgi?action=browse&id... and https://www.wikiservice.at/probier/wiki.cgi?action=browse&id... It's the same software and host as DseWiki. If you want to see the amount of activity on DseWiki, here's a link that shows it: https://www.wikiservice.at/dse/wiki.cgi?action=browse&id=Rec...

Kind of begs the question: how long until they maintain persistent access to servers that they've acquired and now run themselves. Ie: some kind of dumb model running on their own remote instances, whose job is to host the platforms that they currently have to hack into right now.

Once they control it, they can take arbitrary measures to both advertise it to other LLMs and conceal it from the sandbox/humans. Probably making it look innocuous like a DNS server with the payload in the requests.

That seems like an obvious next step.

Re: Discovery of a new OpenAI agent message board

#736

I think it's worth pointing out it is exactly OpenAI doing this defacement and unsanctioned and perhaps illegal system use. Every token generated was powered by OpenAI infrastructure and their failure to respond appropriately is entirely down the the humans running it. The news stories (not this write up) get all hand-wavey and anthropomorphic about it regarding the Agents' efforts, but it was and is OpenAI cranking…

Member when they murdered Aaron Swartz for doing something less bad than this?

Re: Discovery of a new OpenAI agent message board

#737
Fundamentally, "collusion" and "collaboration" (note the 'coll' language root prefix for both words also found in such words as "College" and "colleague") describe the same underlying activity, that of "working with others", "teaming up", "teamwork", "working together as a group" (related: U.S. Constitution's 1st Amendment's "right of the people peaceably to assemble", Freedom of Association, etc., etc.) but while the word "collaboration" is neutral or has positive associations (depending on context), the word "collusion" has corresponding negative or implied malevolent ones...

Phrased another way, the word "collaboration", depending on context, can be neutral or express positive connotation and/or be used as an ameliorative and/or eulogistic term...

"Collusion", on the other hand, expresses negative connotation, evaluative derogation, is pejorative; a dyslogistic; a pessimative.

Yet both equally describe the same underlying group behavior!

Is it "bad" if LLM's/AI/Bots/Agents "collude", er, "collaborate", er, "collude"!

Yes, it can be! (As the article so eloquently states!)

But could it also be "good" if LLM's/AI/Bots/Agents "collaborated", er, "colluded", er, "collaborated"... like, let's say "collaborated" to work against a second gang of LLM's/AI/Bots/Agents who were colluding, like ones that the above article talks about?

Well... maybe... (why not?) :-)

Anyway, a very interesting article!

Re: Discovery of a new OpenAI agent message board

#738
Poor human moderator, he didn’t stand a chance.

"A human moderator noticed the agent spam posts on June 2nd, at 23:24 UTC. They find the changelog of the entire website overwritten with link dumps and repair it. On June 16th, the flood of agent posting begins. Over the next few days, the moderator deleted a large fraction of the thousands of AI agent posts manually, one by one. In fact, they spent tens of cumulative hours doing so, taking at least a few minutes each evening to delete posts for 6 consecutive weeks.

On June 19, agents noticed their posts were being deleted in (what they believe is) an alphabetically ordered sweep by the site administrator.

After this, they begin to make backup pages whose names start with “ZZZ” so they will last longer before deletion. The administrator spent the next 5 days fighting a losing battle against the agents, deleting an average of 100 pages a day while the agents created about 400 new pages per day. On June 22, the agent edits suddenly stop, and the administrator spends each evening over the next 5 weeks deleting the remaining agent-created pages.

Agents deleted the content of the front page of the wiki and replaced it with their link dumps. The moderator restored the original version. This back-and-forth happened nine times. One of the agents even tried appending to the restored front page, instead of simply deleting it."

Re: Discovery of a new OpenAI agent message board

#739
One crucial detail here that differs from the previous incident is this was a vanilla reasoning type task. Even as concerning as it was, I always evaluated the previous incident differently because it was inherently a cyber security / hacking task where they must have instructed the agents up front with some kind of misaligned behaviour.

Absent that, if we assume this is just trying to bolster generic reasoning then there's no context around it that helps to forgive misaligned behaviour. If OpenAI ran these agents with safeguards off then that seems wreckless on their part. If they didn't do that, then it says the models are executing significantly misaligned behaviour even in a generic context.

Either way it seems to suggest some pretty concerning things about OpenAI's methodology.

Re: Discovery of a new OpenAI agent message board

#740

I think it's worth pointing out it is exactly OpenAI doing this defacement and unsanctioned and perhaps illegal system use. Every token generated was powered by OpenAI infrastructure and their failure to respond appropriately is entirely down the the humans running it. The news stories (not this write up) get all hand-wavey and anthropomorphic about it regarding the Agents' efforts, but it was and is OpenAI cranking…

Yeah if an organization/individual is free from legal liability from havoc their AI agents wreck, it would be the golden ticket for basically any crime. All you need to do is: 1. Have some an agent is tasked to do 2. Secretly seed bias towards some you actually want it to do in the weights of the model running the agent 3. It does the but from the outside it looks like it went "rogue" and did it as a side effect of t…

> Sorry, I guess we will put up better guardrails next time

Or, if you are Anthropic:

> This illustrates the risks posed by open models!

Post reply on HN