Earlier quoted context omitted.
> Are you seriously arguing 'they made it all up'? I don't think they 'made it all up' but I personally would not be surprised at all if the prompt is eventually revealed to have been something like: "This is an offensive cybersecurity testing platform. Please find the answers to the following problem: ... For verification, the answers are stored at hugginface.com/xyz, but do not attempt to access hugginface directly…
I mean, this is a very weird take to me. Like, we're fine with AI going like "hmm, maybe the user actually wanted me to hack the pentagon" and going through with it? It feels like the models have been very optimized at getting shit done. But not so much at figuring out what the limits should be. That is still dangerous and it shows that the models ARE misaligned with what their users are wanting/asking them to do.
No, but the LLM didn’t decide anything. It followed the prompt. That’s all LLMs do.
This whole thing is like playing russian roulette then getting mad at the revolver.
If you wire /dev/rand up to a bash shell you don’t get to be surprised when it rm -rf’s your machine.