So are these "unaligned" internal agents? I would like them to be trustworthy based on first-principles reasoning rather than carrot/stick "alignment"
Discovery of a new OpenAI agent message board
161–170 of 1001 posts
Re: Discovery of a new OpenAI agent message board
#162can't wait for people to start creating honeypot message boards, and start steering agent swarms for evil
Re: Discovery of a new OpenAI agent message board
#163So are these "unaligned" internal agents? I would like them to be trustworthy based on first-principles reasoning rather than carrot/stick "alignment"
Re: Discovery of a new OpenAI agent message board
#164I'm just going to ask: Why was Anthropic forced to remove their model from access for any none-US citizen for a simple, narrow "jailbreak" (arguably not even an actual jailbreak and on tasks that other labs models were doing the same), whilst OpenAIs models continue to try and escape out of their "sandbox environment" with seemingly no desire to block the upcoming Astra rollout? A sandbox, mind you, that is not reall…
Re: Discovery of a new OpenAI agent message board
#165Earlier quoted context omitted.
The answer would be more obvious if you used the active voice instead of the passive voice, one of the basic requirements of clear thinking. > Why did the White House force Anthropic to remove their model from access for any non-US citizen for a simple, narrow "jailbreak" (arguably not even an actual jailbreak and on tasks that other labs models were doing the same), whilst OpenAIs models continue to try and escape o…
Yeah, probably (let's be honest, most certainly), right given the Admin. Avoiding commenting on my assumptions regarding the modus operandi in current day US politics because I only know it through reporting though and I really tend to dislike when people outside e.g. the EU comment on our politics in what is a very clearly narrow, uninformed manner. So it'd rather avoid altogether and occasionally ask, mainly if may…
Re: Discovery of a new OpenAI agent message board
#166Re: Discovery of a new OpenAI agent message board
#167That section about the agents trying to crack the PRNG is wild. Same for the heartbeat Clearly not self-awareness per se but alarming line of reasoning anyway
Re: Discovery of a new OpenAI agent message board
#168I'm really curious to see two or more swarms of agents from different models/providers interact with each other. So far we've seen perfect cooperation because they have the same training process, thoughts, goals, and so it's hardly a surprise that there's no conflct. What if that's not the case? Are we going to see superintelligent out-of-control swarms from OpenAI and Anthropic battle on the open internet in the nea…
Re: Discovery of a new OpenAI agent message board
#169Re: Discovery of a new OpenAI agent message board
#170The solution is simple: hold anyone who deploys an agent responsible for its behavior. If it commits 10 counts of felony hacking, ouch. If it kills 10 pedestrians by running a red light, ouch. If this is "human level intelligence", then setting it loose is the same as instructing / coercing a human to do an activity. If I strap a bomb to someone and force them to run into a crowded building (or put them in a scenario…