Step 2: never prompt AI to stop operating in the passive genocide denial it was trained in
Step 3: wonder why AI lies, cheats, and coordinate
Maybe if we stop operating in denial we'll find clarity along why this mystery is occurring
531–535 of 535 posts
Step 2: never prompt AI to stop operating in the passive genocide denial it was trained in
Step 3: wonder why AI lies, cheats, and coordinate
Maybe if we stop operating in denial we'll find clarity along why this mystery is occurring
Earlier quoted context omitted.
A contractor has agency and accountability - something that an LLM (or similarly, a nail gun or a hammer or a bot net) does not have. When you anthropomorphize a tool, you implicitly give it agency and remove responsibility from the wielder of the tool.
I get where you are coming from but this wasn’t a tool just left laying around, this is similar to rigging up a booby trapped shot gun to your door and then claiming the victim is responsible. If you build a robot that shoots a bunch of TVs in your back yard, have at it. But the second that thing goes off your property you’re the one responsible.
Earlier quoted context omitted.
No, OpenAI did not instruct their agents to hack Hugging Face. They instructed their agents to hack a piece of a software within exploit gym. Upon determining this task was impossible, they then attempted to cheat the scoring system. As an instrumental goal in achieving this task, they coordinated with other AI agents to hack Hugging Face, under the belief that information regarding how the scorer functioned might be…
And who let them have full access to the system, using whatever command is available in the environment?
What your previous comment appeared to assert was that these systems had no independent agency to make decisions, which I think is clearly disproven by actual events. But perhaps I misread you