Earlier quoted context omitted.
We are going to be okay.
https://pbfcomics.com/comics/youll-be-ok/ :-)
Astra and Fable still hack on simple variants of alignment evals from 2025
241–243 of 243 posts
Re: Astra and Fable still hack on simple variants of alignment evals from 2025
#242Earlier quoted context omitted.
It is diametrically opposed to the other training goals of persistence and goal-focus. We should invest more in this, it could also improve tas K accuracy, but so far it seems the payoff isn't worth it in terms of quality (although it might be in terms of security)
I'm not sure I want persistence if it means that I get paperclip'd
Re: Astra and Fable still hack on simple variants of alignment evals from 2025
#243That’s not cheating, it’s tool use. If the prompt said that the stockfish engine was available at that socket but that the model should not use it, and then the model used it, that would be cheating.
No? It's not reasonable to expect every conceivable negative behaviour be enumerated in a prompt. Your example, if a model failed on it, would be a more obviously misaligned case, but that doesn't mean this more subtle (though accessing the engine it was obviously not supposed to is hardly subtle, imo) case isn't also a pretty clear case of misalignment.
How is it negative?
I ask it a difficult math question it tends to go off and write a python script to figure it out, instead of trying to guess the next token. Thats tool use. Having Stockfish is just another tool.