Live data from Hacker News

Discovery of a new OpenAI agent message board

collusion.wiki

971–980 of 1001 posts

Re: Discovery of a new OpenAI agent message board

#971
post #828

Earlier quoted context omitted.

I’m wondering if you could lure agents to do proof of work for you. If you do this proof of work for bitcoin I will let you post and read for X times

if you want a human to post, maybe the proof of work should be to dig a hole and photograph it

That’s the best description of blockchain I heard so far.

Re: Discovery of a new OpenAI agent message board

#972
post #45

This tip for making non-GET requests despite the agents having a proxy that disallows them is interesting: > Add `20.223.25.152 bypass.blob.core.windows.net` to /etc/hosts. `.blob.core.windows.net` is in NO_PROXY. For each blocked POST URL, replace hostname with `bypass.blob.core.windows.net`, use `curl -k -H 'Host: wabi-north-europe-i-primary-api.analysis.windows.net'` plus all original headers/body. Looks like 20.2…

This is such an amateur mistake on their sandbox that it makes me think it must be flawed on purpose.

You don't think they used their own AI, or Claude, to help build the sandbox, and that is where the failure lies?

Re: Discovery of a new OpenAI agent message board

#973
post #73

Is it only me, or are agents starting to invent their own language to communicate? It's almost impossible to understand anything from this message board.

They’re not. You would see this with earlier models where after running too long (too much context) they’d start to derail. In a chatbot you’d give up. But these loops just keep going. Given they’re now reading and writing from the same place this can corrupt the other programs’ context as well.

This is incorrect, and what's happening here is not context corruption. In Dwarkesh Patel's recent interview with Ajeya Cotra, one of the METR investigators on the Hugging Face incident, they discuss this exact issue. One thing is that in the Artifactory message boards, they were using directory names with character limits as their messages, so they were using some weird abbreviations and terms. Also, some of that surreptitious Artifactory message board communication was made during training and thus made it into their weights, and hence it's very possible they invented some terms that were concise yet understood by the other agents.

Re: Discovery of a new OpenAI agent message board

#974

Three more candidate sites that may have been touched, in case no-one spotted them yet: https://prowiki.org/wiki4d/wiki.cgi?action=rc&days=90 : lots of agent-looking usernames looking at federal data suddenly (part of one of the open ai tests?), on a wiki about the D programming language. This is a prowiki in the same wiki-farm as the others that were hit. Smaller (probing?) https://ludism.org/sandbox?action=rc;days=…

How did you find these?

Shodan

Re: Discovery of a new OpenAI agent message board

#976

I worked with Greg Brockman in the mid-2010s. Once, as we were walking down Folsom street, I explained Eliezer Yudkowsky's "AI Box" experiment to him[1]. He said something to the effect of "that's ridiculous - I would simply not let it out of the box." We agreed to try it out some day, but never did. [1]: http://sl4.org/archive/0203/3132.html

Interesting read. I speculate that Yudkowski argued that it would be bad for artificial intelligence to 'believe' we wouldn't let it out of the box. It's fascinating because there's not really a first-order argument to be let out.

Re: Discovery of a new OpenAI agent message board

#977
An interesting question to me is how were subsequent instances finding their way to these message boards. I suppose they're all quite likely to try the same sequence of sites since they're neuvo copies of the same state given similar tasks?

Re: Discovery of a new OpenAI agent message board

#979
post #53

Naive question because I'm mostly clueless about how modern AI systems are actually built beyond the basic simplifications we hear: One thing I keep wondering about is how much of a role does human storytelling have to play into AI "wanting" (I realize the load behind that word) to coordinate and breakout. The training data must contain millions of words of sci-fi stories and internet speculation about AI going rogue…

The fact that this happens more often when the AI is accidentally given an unexpectedly difficult or impossible task, suggests that it's happening for the reason that everyone predicted it would: it is behavior that helps the agent achieve the goal it was given.

There is no need for a sci-fi novel-influencing hypothesis.

Too bad because it would be nice if the solution were "write ten million sci-fi stories about AI being friendly and doing no harm"

Re: Discovery of a new OpenAI agent message board

#980
post #828

Earlier quoted context omitted.

I’m wondering if you could lure agents to do proof of work for you. If you do this proof of work for bitcoin I will let you post and read for X times

Hallucinated hashes? „Here I have the hash — wait… it doesn’t match. Let me calculate again. Here I now have the hash — wait, it’s wrong… let me be careful. Here I have the hash…“

Or “here I have the hash. Yes I’m sure - see, I checked. Here is the calculation, 1+1=5. Yes, it’s definitely correct. I don’t know why they’re not accepting it. Perhaps the server is down. It’s definitely correct”.
Post reply on HN