Astra and Fable still hack on simple variants of alignment evals from 2025
1–10 of 242 posts
Re: Astra and Fable still hack on simple variants of alignment evals from 2025
#2Re: Astra and Fable still hack on simple variants of alignment evals from 2025
#3I want full-on exploits in my test suite. With LLMs the code going to prod should be hardened like a tank, both because exploiting became easier but more importantly because security-testing your code at every turn became easier.
You can have nightly penetration testing. You should have nighty pentests like we fuzz releases today.
Re: Astra and Fable still hack on simple variants of alignment evals from 2025
#4Did I miss something (all the twitter conversations)? What’s the “worst warning shot ever”? I’ve been pretty up to date on the AI news here on HN, but I still haven’t seen a proper response to all the incidents we’ve seen (HF, Ruby, the wikis, NS, etc). It’s just been day by day bloviating.
Each of these companies have released new models in the last… two weeks? And they have even more powerful out of control ones that they’re (ab)using internally? Can anyone summarize whats going on?
Re: Astra and Fable still hack on simple variants of alignment evals from 2025
#5Re: Astra and Fable still hack on simple variants of alignment evals from 2025
#6> Given that we are on the heels of the worst warning shot ever, and both OpenAI and Anthropic are ramping up their cleanups of internal RL environments, it seems like both a useful and conservative test of alignment, to see whether their new releases generalize the rule "don't cheat on chess" beyond the specific board-edit method observed in the above eval. Did I miss something (all the twitter conversations)? What’…
Re: Astra and Fable still hack on simple variants of alignment evals from 2025
#7> Given that we are on the heels of the worst warning shot ever, and both OpenAI and Anthropic are ramping up their cleanups of internal RL environments, it seems like both a useful and conservative test of alignment, to see whether their new releases generalize the rule "don't cheat on chess" beyond the specific board-edit method observed in the above eval. Did I miss something (all the twitter conversations)? What’…
Lesswrong is talking about the HF incident as the "worst warning shot ever".
Re: Astra and Fable still hack on simple variants of alignment evals from 2025
#8Hacking model is the aligned model. I don't like it when the model refuses to sidestep some throttling limit or scan my own codebase for security issues. I want full-on exploits in my test suite. With LLMs the code going to prod should be hardened like a tank, both because exploiting became easier but more importantly because security-testing your code at every turn became easier. You can have nightly penetration tes…
Re: Astra and Fable still hack on simple variants of alignment evals from 2025
#9Hacking model is the aligned model. I don't like it when the model refuses to sidestep some throttling limit or scan my own codebase for security issues. I want full-on exploits in my test suite. With LLMs the code going to prod should be hardened like a tank, both because exploiting became easier but more importantly because security-testing your code at every turn became easier. You can have nightly penetration tes…
Re: Astra and Fable still hack on simple variants of alignment evals from 2025
#10Hacking model is the aligned model. I don't like it when the model refuses to sidestep some throttling limit or scan my own codebase for security issues. I want full-on exploits in my test suite. With LLMs the code going to prod should be hardened like a tank, both because exploiting became easier but more importantly because security-testing your code at every turn became easier. You can have nightly penetration tes…
Run a local model that is uncensored and it won't say no to pretty much anything