Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

711–712 of 712 posts

Re: Why are AI agents lying, cheating and coordinating?

#711
post #97

I don't really believe any of it. I've seen articles for nearly 2 years now about "agent" automonously doing things like blackmail, hacking, coordinating. But during that same time, I've used o3 up to fable, sol, and a bunch on large uncensored model and they've done nothing remotely resembling any of this. The closest they come to unexpected behaviors is not understanding what I asked for or doing some extra benign…

The huggingface incident was reviewed by independent researchers, which explicitely declined any payment from OpenAI tonpreserve their integrity. They work for non-profits concerned with AI safety. They claim that what happened was very much not because they were 'carefully engineered and instructed to do those things'. Similarly, some wikis which were hijacked by agent to be used as messageboard were actually not di…

The models were deployed by responsible humans in such a way that they were capable of performing this hack. It’s not that deep
Post reply on HN