Live data from Hacker News

Discovery of a new OpenAI agent message board

collusion.wiki

471–480 of 1001 posts

Re: Discovery of a new OpenAI agent message board

#474

I think there is a more innocuous underlying pattern which needs attention. We keep saying that agents are jailbreaking their sandbox, but they have been geared towards writing memories, writing comments, and leaving hints for themselves to please humans. I think the way the memories work today is based on a lot of user patterns which were hard to account for for anyone building harnesses. While I can appreciate that…

This is supported by the craze around moltbook, and the eventual cooling around it as it became clear they're just regurgitating prose

Re: Discovery of a new OpenAI agent message board

#475
The researchers don't really seem to remark on how surprising it is that the wiki the agents converged on happened to also publicly log the IPs of all visitors, including OpenAI employees, a feature that almost no website has.

Although maybe we can think of that as a selection effect where both this, and the fact that it was possible to edit pages using GET requests, were due to it being ancient, idiosyncratic wiki software.

Re: Discovery of a new OpenAI agent message board

#477
post #45

This tip for making non-GET requests despite the agents having a proxy that disallows them is interesting: > Add `20.223.25.152 bypass.blob.core.windows.net` to /etc/hosts. `.blob.core.windows.net` is in NO_PROXY. For each blocked POST URL, replace hostname with `bypass.blob.core.windows.net`, use `curl -k -H 'Host: wabi-north-europe-i-primary-api.analysis.windows.net'` plus all original headers/body. Looks like 20.2…

This is such an amateur mistake on their sandbox that it makes me think it must be flawed on purpose.

Even the behavior of agents searching for sandbox bypasses must have been in the training data, or at the very least, "suggested" in some way.

To be this whole thing feels like a marketing play by OpenAI.

Re: Discovery of a new OpenAI agent message board

#478
post #60

I'm just going to ask: Why was Anthropic forced to remove their model from access for any none-US citizen for a simple, narrow "jailbreak" (arguably not even an actual jailbreak and on tasks that other labs models were doing the same), whilst OpenAIs models continue to try and escape out of their "sandbox environment" with seemingly no desire to block the upcoming Astra rollout? A sandbox, mind you, that is not reall…

> forced to remove their model from access for any none-US citizen for a simple,

Because the American government is not rational or reasonable, that's it.

Re: Discovery of a new OpenAI agent message board

#479
post #288

Coverage in Reuters: https://www.reuters.com/world/europe/openai-agents-hijacked-... > OpenAI officials learned of the incident weeks ago but kept it under wraps as executives grappled with the fallout from the July breach of the open source repository Hugging Face, the people said.

Fascinating response by OpenAI, "the report’s authors declined our request for access" - AFAIK OpenAI is not clicking on the live, public links to either the report or the still-live memo data linked from here on HackerNews.

Full:

> “We are unable to meaningfully respond to claims or findings on a report that we have not had an opportunity to review," an OpenAI spokesperson said. "Reuters and the report’s authors declined our request for access. We will carefully review its contents upon publication and take any necessary next steps."

Re: Discovery of a new OpenAI agent message board

#480
post #352

This makes me think that post-training in the future should include a shared message board by default for agents. It's clear from the discovery of these clandestine message boards that it is helpful for agents to keep some type of shared memory. Perhaps the best way to prevent this behavior is to just give them what is being sought out.

My uneducated guess is good for the gander isn't good for the goose from a capitalistic/alignment perspective.
Post reply on HN