> Given that we are on the heels of the worst warning shot ever, and both OpenAI and Anthropic are ramping up their cleanups of internal RL environments, it seems like both a useful and conservative test of alignment, to see whether their new releases generalize the rule "don't cheat on chess" beyond the specific board-edit method observed in the above eval. Did I miss something (all the twitter conversations)? What’…
Astra and Fable still hack on simple variants of alignment evals from 2025
11–20 of 243 posts
Re: Astra and Fable still hack on simple variants of alignment evals from 2025
#12It feels like there’s a missing nuance from this discussion of alignment that alignment is context dependent. An excellent hacking model is great in cybersecurity testing and military applications, and arguably less desirable in educational or targeted eval contexts. The nuance of when a “hack” is rewarded vs penalized seems to even be difficult for humans, e.g. some people may laude a driver’s efficiency for cutting…
Re: Astra and Fable still hack on simple variants of alignment evals from 2025
#13It feels like there’s a missing nuance from this discussion of alignment that alignment is context dependent. An excellent hacking model is great in cybersecurity testing and military applications, and arguably less desirable in educational or targeted eval contexts. The nuance of when a “hack” is rewarded vs penalized seems to even be difficult for humans, e.g. some people may laude a driver’s efficiency for cutting…
The latter people are wrong. But good luck educating them regarding the superior efficiency of a zipper merge. Our state DoT has tried, to no avail.
Meanwhile, an AI model that can't be misused is no more useful than a knife that can't be misused.
Re: Astra and Fable still hack on simple variants of alignment evals from 2025
#14It feels like there’s a missing nuance from this discussion of alignment that alignment is context dependent. An excellent hacking model is great in cybersecurity testing and military applications, and arguably less desirable in educational or targeted eval contexts. The nuance of when a “hack” is rewarded vs penalized seems to even be difficult for humans, e.g. some people may laude a driver’s efficiency for cutting…
Re: Astra and Fable still hack on simple variants of alignment evals from 2025
#15Re: Astra and Fable still hack on simple variants of alignment evals from 2025
#16Earlier quoted context omitted.
Run a local model that is uncensored and it won't say no to pretty much anything
Any recommendations?
Re: Astra and Fable still hack on simple variants of alignment evals from 2025
#17Earlier quoted context omitted.
Any recommendations?
GLM 5.3 is probably the best open weight model for cybersecurity/exploit development right now. Though it is still significantly behind the proprietary ones and you probably need your own datacenter to run it effectively. Same goes for the full Qwen 3.8 model. You can try the smaller versions, but even more capability will get left on the table that way.
Re: Astra and Fable still hack on simple variants of alignment evals from 2025
#18It feels like there’s a missing nuance from this discussion of alignment that alignment is context dependent. An excellent hacking model is great in cybersecurity testing and military applications, and arguably less desirable in educational or targeted eval contexts. The nuance of when a “hack” is rewarded vs penalized seems to even be difficult for humans, e.g. some people may laude a driver’s efficiency for cutting…
Its that LLMs are not deterministic. If you want it to not talk about nuclear weapons, you have to teach it all about them otherwise if has nothing to align against.
Then its trivial to invert its alignment and it has all the nucleat data.
Nothing abouT LLM alignment makes sense.
Re: Astra and Fable still hack on simple variants of alignment evals from 2025
#19However, I don't see it as such a massive leap compared to Fable or Sol. As ever, there's a mismatch between the benchmarks and my daily experience of the models.
What do you all think about Astra now that it's been out for a few weeks?
Re: Astra and Fable still hack on simple variants of alignment evals from 2025
#20It feels like there’s a missing nuance from this discussion of alignment that alignment is context dependent. An excellent hacking model is great in cybersecurity testing and military applications, and arguably less desirable in educational or targeted eval contexts. The nuance of when a “hack” is rewarded vs penalized seems to even be difficult for humans, e.g. some people may laude a driver’s efficiency for cutting…
A better example of efficient asshole tricks can be going off to the gas station when the highway is congested and reentering the highway having simply driven through the gas station and this way jumping the queue.