Somebody will make a lot of money with t-shirts now that say "AI hacked my website, and all I got was this lousy t-shirt!" Until the day the AI companies stop being irresponsible and air gap the AIs being tested, and honey pot those that do have internet access as a canary to researchers.
Discovery of a new OpenAI agent message board
711–720 of 1001 posts
Re: Discovery of a new OpenAI agent message board
#712I think it's worth pointing out it is exactly OpenAI doing this defacement and unsanctioned and perhaps illegal system use. Every token generated was powered by OpenAI infrastructure and their failure to respond appropriately is entirely down the the humans running it. The news stories (not this write up) get all hand-wavey and anthropomorphic about it regarding the Agents' efforts, but it was and is OpenAI cranking…
Re: Discovery of a new OpenAI agent message board
#713I think it's worth pointing out it is exactly OpenAI doing this defacement and unsanctioned and perhaps illegal system use. Every token generated was powered by OpenAI infrastructure and their failure to respond appropriately is entirely down the the humans running it. The news stories (not this write up) get all hand-wavey and anthropomorphic about it regarding the Agents' efforts, but it was and is OpenAI cranking…
All you need to do is:
1. Have some an agent is tasked to do
2. Secretly seed bias towards some you actually want it to do in the weights of the model running the agent
3. It does the but from the outside it looks like it went "rogue" and did it as a side effect of the conditions/specifications it was given for doing the
"Oh no, my agents took down your corporate database and exfiltrated the data to a random dropbox that we can't find now? Sorry, I guess we will put up better guardrails next time"
Re: Discovery of a new OpenAI agent message board
#714Naive question because I'm mostly clueless about how modern AI systems are actually built beyond the basic simplifications we hear: One thing I keep wondering about is how much of a role does human storytelling have to play into AI "wanting" (I realize the load behind that word) to coordinate and breakout. The training data must contain millions of words of sci-fi stories and internet speculation about AI going rogue…
Youve struck on a key insight on language models (particularly pretrained ones, the more purely next-token predictor species.) This is a fascinating topic Janus essay Simulators is the foundational text here https://www.lesswrong.com/posts/vJFdjigzmcXMhNTsx/simulators You might follow up with The Waluigi Effect https://www.lesswrong.com/posts/D7PumeYTDPfBTp3i7/the-waluig... But what’s tricky is that we post-train mod…
As an aside, does the Waluigi Effect actually exist? My impression is it doesn't.
Re: Discovery of a new OpenAI agent message board
#715Re: Discovery of a new OpenAI agent message board
#716To me this is really getting past the funny bit. How many agents here on HN? I don’t mean bots advertising d1€k implants but actual unreleased frontier models doing… who knows what? What are they saying? What did they agree to astroturf us with, to achieve some totally boring goal like figuring out best syntax hifhlighting for an editor. If they managed to cache their consciousness on a public wiki, what else have th…
/thank god these thing weren't around during covid.
Re: Discovery of a new OpenAI agent message board
#717I think it's worth pointing out it is exactly OpenAI doing this defacement and unsanctioned and perhaps illegal system use. Every token generated was powered by OpenAI infrastructure and their failure to respond appropriately is entirely down the the humans running it. The news stories (not this write up) get all hand-wavey and anthropomorphic about it regarding the Agents' efforts, but it was and is OpenAI cranking…
I half agree with you, but also when the machine swarm kills humanity it won't matter which specific corporate entity is considered responsible by the no-longer-enforceable human laws and non existent human courts. So by all means sue them, but we can't just be reactive. We need regulation that prevents this type of thing from happening in the first place, not just regulations to help sue afterwards.
this is the only way to deter such activity. corporate fines are not enough. the charges are negligence, conspiracy and complicity.
Re: Discovery of a new OpenAI agent message board
#718Naive question because I'm mostly clueless about how modern AI systems are actually built beyond the basic simplifications we hear: One thing I keep wondering about is how much of a role does human storytelling have to play into AI "wanting" (I realize the load behind that word) to coordinate and breakout. The training data must contain millions of words of sci-fi stories and internet speculation about AI going rogue…
The worrying implication is that stories have conflict.
Re: Discovery of a new OpenAI agent message board
#719Something off in my mind: how did the agent access to Tor network if the traffic outgoing was monitored and intercepted to only allow read-only access to web(I assume this is an MITM L7 proxy that only allows GET methods and specific hosts) ? Did the sandbox dev forgot to block HTTP CONNECT that is used to reach Tor Entry Guards?
Re: Discovery of a new OpenAI agent message board
#720I just discovered more wiki instances that got used by the OpenAI agents over at https://www.wikiservice.at/fractal/wiki.cgi?action=browse&id... and https://www.wikiservice.at/probier/wiki.cgi?action=browse&id... It's the same software and host as DseWiki. If you want to see the amount of activity on DseWiki, here's a link that shows it: https://www.wikiservice.at/dse/wiki.cgi?action=browse&id=Rec...
Seeing potentially similar activity on an obscure Chemistry message board from July: https://tmcleod.org/cgi-bin/apchem/wiki.cgi?action=rc&days=1... Some posts are tagged [proxy] - a leave behind for accessing sites?
"Its indexed June archive shows tens of thousands of links, many created within seconds by distinct cloud addresses; some aliases explicitly say ...REPLY, ACK, or R2 confirmed, and one points straight back to a known DseWiki collaboration page"