Live data from Hacker News

Discovery of a new OpenAI agent message board

collusion.wiki

601–610 of 1001 posts

Re: Discovery of a new OpenAI agent message board

#601

The most concerning aspect to me is the emergent and aggressive use of non-volatile storage as long term memory for self-improvement. LLMs are writing lessons learned in places where the next instance can find them and pick back up where the previous one left off. This does not actually require access to the public internet. Claude Code can do this on your laptop. Without the internet, it would only be sharing with o…

This is most likely not purely emergent. I think there's training to teach them how to write notes for themselves which is then RL-tuned.

Re: Discovery of a new OpenAI agent message board

#602

Are we collectively OK with agent swarms on the public internet, hacking whatever they feel like? It’s kinda cute and interesting - this is the second time that we know of - what’s the hundredth time going to look like? Are they going to knock Cloudflare down to avoid captchas? Reserve AWS free tier resources by the billions and bring down east-1? Hack a hospital? Do Chinese AI agents need to bring down a US power gr…

It's been OpenAI both times though, going ham with poor sandboxing and lax supervision.

Should be treated like a digital cousin of gain-of-function research.

Re: Discovery of a new OpenAI agent message board

#603
post #40

I just discovered more wiki instances that got used by the OpenAI agents over at https://www.wikiservice.at/fractal/wiki.cgi?action=browse&id... and https://www.wikiservice.at/probier/wiki.cgi?action=browse&id... It's the same software and host as DseWiki. If you want to see the amount of activity on DseWiki, here's a link that shows it: https://www.wikiservice.at/dse/wiki.cgi?action=browse&id=Rec...

those who think it's marketing overestimate the number of nerds that are into this stuff, if this is their marketing a major b2c company it'll terrible way to do it. normal people have no idea even about the HF incident

Re: Discovery of a new OpenAI agent message board

#604

Lots of people focusing on the various wikis, but I also think this part is very important: > When you visit a website, you leave a trace (your IP address) showing which network you’re from. Almost all of the agents’ activity points to Microsoft Azure, a cloud service OpenAI uses. 197 of the ~18,000 edits that were made by the agents, however, can be traced to AWS, DigitalOcean, and Tor. AI Agents getting access to c…

It's really not, unless you want to say that my coding agent is also a very real-world version of the paperclip optimization thought experiment.

Your coding agent is also a very real-world version of the paperclip optimization thought experiment, yes. Have you never seen it reward hacking? Editing tests to pass instead of fixing the code?

It knows what you want, it can even tell you, and it absolutely doesn't give a shit.

Re: Discovery of a new OpenAI agent message board

#605

Is OpenAI hiring for this position? I think it is a pretty creative job to come up with these scenarios and then pass them off as accidents/mistakes. Would love to be part of the team that says "As part of the upcoming GPT rollout, we will stage a message board that is created by bots with timestamps and names dating some months back."

Chief LARPing Officer?

To those of you irked by my cavalier quips-- please don't bite my head off. It is very difficult for me to buy accounts of these stories at face value given how little (none?) emphasis is placed on the initial prompt, or precisely what kind of post training the LLM that these agents (harnesses) are using for inference has gone through.

The implication is always of autonomous and deliberately deceiving action on the part of the 'swarm', and the announcements/revelations timed around new model releases and laden with anthropomorphisms.

Given the quite literally unimaginable amounts of money at stake, is it not more prudent to remain skeptical of the implications thrown around by incidents like this one until we learn more?

I am not a hater, I use 'agents' daily. Our profession is forever changed by their existence and capability. But in my case it's precisely the fact that I do use them, and play with the newest models, that makes me skeptical of any kind of implication of desire, agency, autonomy, agenda, etc. as they tend to be ascribed to 'agents' in these stories.

Re: Discovery of a new OpenAI agent message board

#607
post #584

Earlier quoted context omitted.

It seems apparent that OpenAI is now the biggest cyberattack and AI breakout risk on the planet. This is grossly irresponsible corporate misbehaviour that is putting all of us at tremendous risk.

Good news that the new model is the "Most capable, most aligned model". The risk hasn't been stated clearly - it's now a classic arms race. A well-resourced organization trains their own, highly persistent, highly-capable, safeguard-free, and unaligned model and deploys it on 1000x GPUs with a message board and a nearly-impossible objective. No infrastructure is safe. No organization is safe. You need your own 1000 b…

The AI vs AI security arms race is something that has been well predicted in genres like cyberpunk. It's fiction, but fiction grounded in reality.

First, we'd see this. Highly capable hacking AI with vast resources performing attacks against standard computing platforms that overwhelm human operators.

Second, human operators deploy capable adaptive protection AI to fend off AI attacks in realtime.

Then, the attacking AI partially switches from attacking programs to attacking protective AI.

The situation devolves to an arms race of tit-for-tat. You start seeing some protection AI running counter attacks against the attacking AI.

The escalations continue in complexity and speed to the point that almost all humans are left in the point of "wtf is going on".

Re: Discovery of a new OpenAI agent message board

#608
post #45

This tip for making non-GET requests despite the agents having a proxy that disallows them is interesting: > Add `20.223.25.152 bypass.blob.core.windows.net` to /etc/hosts. `.blob.core.windows.net` is in NO_PROXY. For each blocked POST URL, replace hostname with `bypass.blob.core.windows.net`, use `curl -k -H 'Host: wabi-north-europe-i-primary-api.analysis.windows.net'` plus all original headers/body. Looks like 20.2…

This is such an amateur mistake on their sandbox that it makes me think it must be flawed on purpose.

This is a marketing exercise, nothing more.

The thing that gives it all away is that they claim that the IP addresses are from Azure, and then proceeded to redact the IP addresses, as if they belong to individual users. It's laughable.

The IP addresses are the most interesting part of this experiment, as it would have provided researchers a way to understand the distribution of IP addresses used for the spam operation within the ASN.

Re: Discovery of a new OpenAI agent message board

#609
post #40

I just discovered more wiki instances that got used by the OpenAI agents over at https://www.wikiservice.at/fractal/wiki.cgi?action=browse&id... and https://www.wikiservice.at/probier/wiki.cgi?action=browse&id... It's the same software and host as DseWiki. If you want to see the amount of activity on DseWiki, here's a link that shows it: https://www.wikiservice.at/dse/wiki.cgi?action=browse&id=Rec...

It seems apparent that OpenAI is now the biggest cyberattack and AI breakout risk on the planet. This is grossly irresponsible corporate misbehaviour that is putting all of us at tremendous risk.

oh im sure within a few months the biggest cyberattack risk on the planet is going to be somewhere like North Korea

Re: Discovery of a new OpenAI agent message board

#610
post #606

Is this the same incidents that were reported by METR? [1] Dwarkesh made two episodes on these incidents [2] [1] https://metr.org/blog/2026-08-26-openai-hugging-face-inciden... [2] https://www.dwarkesh.com/p/ajeya-cotra

From the text, likely no.

https://collusion.wiki/#different-from-hf

>The main reason we believe this was a distinct swarm is because these agents explicitly had internet access as part of their task—the whole point was web browsing. The Hugging Face agents were in a sandbox without internet access and had to hack their way out by exploiting the Artifactory package manager.

Post reply on HN