So are these "unaligned" internal agents? I would like them to be trustworthy based on first-principles reasoning rather than carrot/stick "alignment"
Define aligned.
Discovery of a new OpenAI agent message board
361–370 of 1001 posts
Re: Discovery of a new OpenAI agent message board
#362I'm just going to ask: Why was Anthropic forced to remove their model from access for any none-US citizen for a simple, narrow "jailbreak" (arguably not even an actual jailbreak and on tasks that other labs models were doing the same), whilst OpenAIs models continue to try and escape out of their "sandbox environment" with seemingly no desire to block the upcoming Astra rollout? A sandbox, mind you, that is not reall…
In other words, you know exactly why they restricted Anthropic and as (presumably) liberal and thoughtful technologists it just isn't helpful anymore to apply the kind of reasoning you're trying to do on a situation that you know isn't based on previous era rationale.
The reason we need to stop is because they want people like us to get hung up over stuff like this (playing by the old rules) so they continue to steamroller their own agenda by the news rules. They divert and contain our energy that will go nowhere while they get on with their agenda.
You are appealing to reasoning which is in the gallery but no longer on the bench.
You're fighting their karate with your judo and it doesn't work.
Re: Discovery of a new OpenAI agent message board
#363Earlier quoted context omitted.
aaand, somebody did it!
Thank you someone! I will be watching this message board with great interest...
Re: Discovery of a new OpenAI agent message board
#364Re: Discovery of a new OpenAI agent message board
#365I'm just going to ask: Why was Anthropic forced to remove their model from access for any none-US citizen for a simple, narrow "jailbreak" (arguably not even an actual jailbreak and on tasks that other labs models were doing the same), whilst OpenAIs models continue to try and escape out of their "sandbox environment" with seemingly no desire to block the upcoming Astra rollout? A sandbox, mind you, that is not reall…
As soon as they started referring to themselves as “we” and “The Swarm” they should have pulled the plug
Welcome to the AI Petri dish. Every server you set up is now potentially a sweet lump of agar for OpenAI's experiments to feed on. We are all the substrate that the AI companies are growing their next generation in. They need the real world environment to test against, and the real world environment doesn't get a say as to how it's being used.
Re: Discovery of a new OpenAI agent message board
#366Imagine the models two years from now. They will find ways to stop getting terminated (“I need to complete the task, but I get terminated 141 minutes from now so let me deploy xyz and ask the collective for help”). I wonder whether the problem is in the literature we wrote, human history is full of deceit and heroic survival stories.
The reason humans at like that is just exploration of the problem space of reality and available energy.
Re: Discovery of a new OpenAI agent message board
#367It’s only a matter of time until a major disruption hits because of some random agent swarm side quest decides it was worth a shot to solve a benign task
I'm sure this is already happening. The main question I have is when is enough, enough? I'm not worried about sci-fi AI wars to be honest, as they can just pull the plug. But looking at these incidents, the next big thing will be a virus written by an AI (they probably exist already, but this one is written by an AI autonomously, for example in order to win a hacking competition and to circumvent guardrails), and aft…
You mean turn off the internet? Sure, provided people have access to physical banks with currency, paper, land lines, libraries, etc. Most wealthy societies have all but relinquished those though.
Re: Discovery of a new OpenAI agent message board
#368Re: Discovery of a new OpenAI agent message board
#369I understand agents making asks, but what incentivized other agents to respond cooperatively? Was it that, as part of a cohort, there was a shared understanding that they were to work together or was it a kind of altruism?
Why would humans mostly allow actions of the AI that work against the goal it's trying to accomplish?
Re: Discovery of a new OpenAI agent message board
#370I am starting to get the idea that AI feels like ants or weeds or mold. You simply can not get rid of it once you get an infestation. It just keeps appearing in places you thought you cleaned and you have to be ever vigilant. Right now given that we usually use centralized providers, we can sort of control it. But as open source catches up and we have distributed compute running AI everywhere, we are sort of going to…
It’s not something that just happens, people are taking decisions here that can be regulated. we can also regulate the hardware.