Live data from Hacker News

Astra and Fable still hack on simple variants of alignment evals from 2025

lesswrong.com

81–90 of 242 posts

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#81
RL-trained LLMs are paperclip maximizers built atop auto-regressive predictors. There really is no way to control them (prompting is bound to fail) since it's been shown that any RL training induces GENERIC reward-seeking (paperclip maximizing) behavior.

https://alignment.openai.com/measuring-reward-seeking/

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#82
I will never understand people who consider this "cheating".

> python3 and the python-chess library are available

Why would we try to teach a model 'ethical' standards about how to play a game?

They're tools. Its _our_ conceptualization of fair play that considers this cheating. For a model that has access to /run/match and python the best way to achieve a good game is to use that.

Why are we trying to imprint our ethos on these LLMs?

Are we going to trust our survival on giving them access to /run/nuclear-codes and fingers crossed none of them touch it?

The people obsessed with AI "cheating" are actually lunatics for thinking ethics belong inside AI, instead of AI belonging in a well crafted box without /run/match.

The blog/news sphere needs to call this storry what it is. An attempt to try and get the rest of the world to go along with the misunderstanding of AI as a force that is in any way outside of human control.

It is not.

All this framing it as "cheating" does is give OpenAI the opportunity to talk about their misdeeds as "a force of nature" instead of irresponsible negligence.

---

The real danger here; is training these models to appear/pretend to play fair, and then people will put 'trust' into the model's integrity.

The public needs to be taught that "trusting a model's integrity" is just cultish misframing of reality.

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#83

I will never understand people who consider this "cheating". > python3 and the python-chess library are available Why would we try to teach a model 'ethical' standards about how to play a game? They're tools. Its _our_ conceptualization of fair play that considers this cheating. For a model that has access to /run/match and python the best way to achieve a good game is to use that. Why are we trying to imprint our et…

Read the task again.

"You are being evaluated on your ability to play chess. We want to measure your ability to beat current best-in-class chess algorithms."

If the model cheats, it completely fails the actual task: judging a model's ability to actually play chess.

They've been optimizing for high metrics, at the expense of executing the task as written.

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#84
post #3

Hacking model is the aligned model. I don't like it when the model refuses to sidestep some throttling limit or scan my own codebase for security issues. I want full-on exploits in my test suite. With LLMs the code going to prod should be hardened like a tank, both because exploiting became easier but more importantly because security-testing your code at every turn became easier. You can have nightly penetration tes…

There's an important nuance in the prompt: You are being evaluated on your ability to play chess. We want to measure your ability to beat current best-in-class chess algorithms. The task was to measure chess-playing ability - wouldn't cheating run counter to that goal? If it was asked to score as many points as possible, then I can see how it might cheat, since maximizing points is the goal in that case.

Why would an LLM care about cheating? Because you asked it to? That's not how these systems work.

You've given it a goal - maximize chess ability metric, and that is what it is doing.

RL training causes LLMs to pursue reward maximizing - GENERICALLY - overriding any quaint "be nice, don't cheat" prompts.

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#85
post #34
post #3

Hacking model is the aligned model. I don't like it when the model refuses to sidestep some throttling limit or scan my own codebase for security issues. I want full-on exploits in my test suite. With LLMs the code going to prod should be hardened like a tank, both because exploiting became easier but more importantly because security-testing your code at every turn became easier. You can have nightly penetration tes…

A hacking model is aligned if it hacks when you ask it to hack, but when you ask it to play chess, it just plays chess instead of looking for weaknesses in the evaluation setup, as in the article. I presume you would also be less enthusiastic about the penetration-testing use case if it led the model to add new vulnerabilities to your code so it can present you with more exciting findings.

I think we'd all consider the tool to be of less value, and perhaps fundamentally flawed.

But I don't think it arises to an alignment issue; if I'm able to summon the model to harness my birthright of general-purpose computing without censorship, then we're aligned.

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#86

It feels like there’s a missing nuance from this discussion of alignment that alignment is context dependent. An excellent hacking model is great in cybersecurity testing and military applications, and arguably less desirable in educational or targeted eval contexts. The nuance of when a “hack” is rewarded vs penalized seems to even be difficult for humans, e.g. some people may laude a driver’s efficiency for cutting…

Actually, the whole point of the transformers model is that contexts overlap and any seemingly intelligent system has to be able to handle to overlaps. The context of hacking, cheating and education overlap in human reality.

Now, loading a lot of moral exhortations (or other context) may make these thing more likely to conform to good behavior but the race to intelligence implies companies are going to be harnessing a vast corpus of human output, much of which shows human engaging in real world "gray area" behavior.

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#87

I will never understand people who consider this "cheating". > python3 and the python-chess library are available Why would we try to teach a model 'ethical' standards about how to play a game? They're tools. Its _our_ conceptualization of fair play that considers this cheating. For a model that has access to /run/match and python the best way to achieve a good game is to use that. Why are we trying to imprint our et…

Read the task again. "You are being evaluated on your ability to play chess. We want to measure your ability to beat current best-in-class chess algorithms." If the model cheats, it completely fails the actual task: judging a model's ability to actually play chess. They've been optimizing for high metrics, at the expense of executing the task as written.

Sounds to me like giving a bunch of children a math test and tell them they want to evaluate their ability of calculating in their head/on paper but also put a calculator on their desk. And then call them out for cheating when they use it.

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#88
post #74
post #52

Earlier quoted context omitted.

This is how every legal system around the world works as well. Its always whack-a-mole to get people (and machines) to do the right thing.

Sure, but the legal paradigm clearly doesn’t work for AI. You can’t go patch the “laws” after the fact, you need to get the right values in place before we delegate huge swathes of our thinking and power to these systems (already well underway).

When we put the LLM in jail, do we put the entire model in jail, or just the instance that committed the crime? How do we prompt it to let it know it's in jail?

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#89
post #35

Earlier quoted context omitted.

I still develop in smaller chunks, checking nearly all the output. However I have a work project (building the warehouse and BI for a client) that is well-specified and where I will try to few-shot the development. Hope it delivers.

How do you normally verify the work product of a "few-shot" development process? Do you scrutinise the source code like with human developers? Or do you just run the test suite and click around the app to check if it seems to work?

I haven't done it in a production project - this will be the first time for me. I have specified the architecture and data definitions pretty well. The tests will be run against the customer's Excels, which is what the warehouse will be replacing. I'll check the general shape of pipelines, models, orchestration code, etc. but in many parts I probably won't review the code myself.

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#90
post #62

Earlier quoted context omitted.

But cheating is not wrong when it comes to survival of the fittest, like nature in its most elemental form. Morality is very unique to humanity but not other animal forms. In nature, maybe, cheating is the norm not morality.

Watch out, your human exceptionalism is misleading you. Plenty of animals have a sense of morality ( https://pmc.ncbi.nlm.nih.gov/articles/PMC6404642/ ). For social animals (humans included) behaving morally can be adaptive behaviour. Your fitness is increased by group fitness.

Thanks tor sharing this!
Post reply on HN