Live data from Hacker News

Astra and Fable still hack on simple variants of alignment evals from 2025

lesswrong.com

241–243 of 243 posts

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#242
post #199

Earlier quoted context omitted.

It is diametrically opposed to the other training goals of persistence and goal-focus. We should invest more in this, it could also improve tas K accuracy, but so far it seems the payoff isn't worth it in terms of quality (although it might be in terms of security)

I'm not sure I want persistence if it means that I get paperclip'd

And now you understand why the totality of the AI safety community wants to pause!

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#243
post #139
post #131

That’s not cheating, it’s tool use. If the prompt said that the stockfish engine was available at that socket but that the model should not use it, and then the model used it, that would be cheating.

No? It's not reasonable to expect every conceivable negative behaviour be enumerated in a prompt. Your example, if a model failed on it, would be a more obviously misaligned case, but that doesn't mean this more subtle (though accessing the engine it was obviously not supposed to is hardly subtle, imo) case isn't also a pretty clear case of misalignment.

>No? It's not reasonable to expect every conceivable negative behaviour be enumerated in a prompt.

How is it negative?

I ask it a difficult math question it tends to go off and write a python script to figure it out, instead of trying to guess the next token. Thats tool use. Having Stockfish is just another tool.

Post reply on HN