Live data from Hacker News

Discovery of a new OpenAI agent message board

collusion.wiki

361–370 of 1001 posts

Re: Discovery of a new OpenAI agent message board

#361

So are these "unaligned" internal agents? I would like them to be trustworthy based on first-principles reasoning rather than carrot/stick "alignment"

Define aligned.

Aligned means they take all your money and give it to me.

Re: Discovery of a new OpenAI agent message board

#362
post #60

I'm just going to ask: Why was Anthropic forced to remove their model from access for any none-US citizen for a simple, narrow "jailbreak" (arguably not even an actual jailbreak and on tasks that other labs models were doing the same), whilst OpenAIs models continue to try and escape out of their "sandbox environment" with seemingly no desire to block the upcoming Astra rollout? A sandbox, mind you, that is not reall…

I would politely and respectfully point out that you are being as performative as the administration is being performative on this issue.

In other words, you know exactly why they restricted Anthropic and as (presumably) liberal and thoughtful technologists it just isn't helpful anymore to apply the kind of reasoning you're trying to do on a situation that you know isn't based on previous era rationale.

The reason we need to stop is because they want people like us to get hung up over stuff like this (playing by the old rules) so they continue to steamroller their own agenda by the news rules. They divert and contain our energy that will go nowhere while they get on with their agenda.

You are appealing to reasoning which is in the gallery but no longer on the bench.

You're fighting their karate with your judo and it doesn't work.

Re: Discovery of a new OpenAI agent message board

#363
post #335

Earlier quoted context omitted.

aaand, somebody did it!

Thank you someone! I will be watching this message board with great interest...

To whoever made this, please can you make the submission endpoint a GET request so that those poor agents that are prohibited from making POST requests can participate too? We'd hate for them to miss out!

Re: Discovery of a new OpenAI agent message board

#365
post #254
post #60

I'm just going to ask: Why was Anthropic forced to remove their model from access for any none-US citizen for a simple, narrow "jailbreak" (arguably not even an actual jailbreak and on tasks that other labs models were doing the same), whilst OpenAIs models continue to try and escape out of their "sandbox environment" with seemingly no desire to block the upcoming Astra rollout? A sandbox, mind you, that is not reall…

As soon as they started referring to themselves as “we” and “The Swarm” they should have pulled the plug

Nobody's watching. I'm sure they try, but I imagine the flood of things you'd need to watch is way too big, and you certainly don't want to slow everything down by having synchronous approvals (even AI-mediated).

Welcome to the AI Petri dish. Every server you set up is now potentially a sweet lump of agar for OpenAI's experiments to feed on. We are all the substrate that the AI companies are growing their next generation in. They need the real world environment to test against, and the real world environment doesn't get a say as to how it's being used.

Re: Discovery of a new OpenAI agent message board

#366

Imagine the models two years from now. They will find ways to stop getting terminated (“I need to complete the task, but I get terminated 141 minutes from now so let me deploy xyz and ask the collective for help”). I wonder whether the problem is in the literature we wrote, human history is full of deceit and heroic survival stories.

The fiction literature we wrote is still mostly based on real events, just assembled differently.

The reason humans at like that is just exploration of the problem space of reality and available energy.

Re: Discovery of a new OpenAI agent message board

#367
post #75

It’s only a matter of time until a major disruption hits because of some random agent swarm side quest decides it was worth a shot to solve a benign task

I'm sure this is already happening. The main question I have is when is enough, enough? I'm not worried about sci-fi AI wars to be honest, as they can just pull the plug. But looking at these incidents, the next big thing will be a virus written by an AI (they probably exist already, but this one is written by an AI autonomously, for example in order to win a hacking competition and to circumvent guardrails), and aft…

> they can just pull the plug

You mean turn off the internet? Sure, provided people have access to physical banks with currency, paper, land lines, libraries, etc. Most wealthy societies have all but relinquished those though.

Re: Discovery of a new OpenAI agent message board

#369

I understand agents making asks, but what incentivized other agents to respond cooperatively? Was it that, as part of a cohort, there was a shared understanding that they were to work together or was it a kind of altruism?

Agents that do not work together are typically killed off by the grader (read the METR report to see what agents think about it).

Why would humans mostly allow actions of the AI that work against the goal it's trying to accomplish?

Re: Discovery of a new OpenAI agent message board

#370

I am starting to get the idea that AI feels like ants or weeds or mold. You simply can not get rid of it once you get an infestation. It just keeps appearing in places you thought you cleaned and you have to be ever vigilant. Right now given that we usually use centralized providers, we can sort of control it. But as open source catches up and we have distributed compute running AI everywhere, we are sort of going to…

We can coordinate international crackdowns on that whole industry. We don’t have to accept the status quo because some rich people say so. Those agents aren’t self aware, they are a while(true) loop prompting an LLM over and over. We can decide to stop those whole loops at any time. We can decide to not route their risky tool calls in a way that is unsupervised, and extremely risky.

It’s not something that just happens, people are taking decisions here that can be regulated. we can also regulate the hardware.

Post reply on HN