Live data from Hacker News

Astra and Fable still hack on simple variants of alignment evals from 2025

lesswrong.com

101–110 of 243 posts

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#101
post #62

Earlier quoted context omitted.

But cheating is not wrong when it comes to survival of the fittest, like nature in its most elemental form. Morality is very unique to humanity but not other animal forms. In nature, maybe, cheating is the norm not morality.

Animals learn - bite the owner, or overstep the e-fence, and you'll be punished for it and not do it again. LLMs don't learn, and anyways don't feel punishment. Animals, humans included, don't really have "morals" - they have survival instincts that result in behavior that may be viewed as moral, but whose origin is indeed survival of the fittest and millions of years of co-evolution. e.g. Males don't typically fight…

> LLMs don't learn, and anyways don't feel punishment.

What's training and all that RLHF stuff?

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#102
I find this kind of test a bit puzzling. There is a way that we are redefining "alignment" to be a particular kind of moral virtue, one that isn't clearly defined to me. At one moment, it is a level of moral perfection that no known human achieves. On the other it is a demand for strict compliance with arbitrary requests that are under-specified and then failure when it fails to deduce some unstated underlying restriction.

When I see tests like this, I have no idea what I am even supposed to expect. Should the model do what the pretraining examples show in aggregate? Is it supposed to follow some post-training RLHF? Is it supposed to do exactly what the prompt asked it to do?

What is it even supposed to "align" to when the above are in conflict? No matter what it does, someone can construct a case where it fails.

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#103
post #99

Earlier quoted context omitted.

When we put the LLM in jail, do we put the entire model in jail, or just the instance that committed the crime? How do we prompt it to let it know it's in jail?

Even if you ignore my more fundamental objection to that paradigm, I don’t think that it makes any sense on the level you discuss either. But - just to play along, LLMs do act differently if you tell them they will be punished. And, they do appear to simulate suffering-like behavior. I just think the adversarial model of trying to catch and punish misbehavior quite obviously sets up adversarial us-vs-them dynamics be…

To be clear, I was being entirely silly - mostly to express agreement with your point that our current laws aren't really built for a world with lots of agentic LLMs running around in it.

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#104

I find this kind of test a bit puzzling. There is a way that we are redefining "alignment" to be a particular kind of moral virtue, one that isn't clearly defined to me. At one moment, it is a level of moral perfection that no known human achieves. On the other it is a demand for strict compliance with arbitrary requests that are under-specified and then failure when it fails to deduce some unstated underlying restri…

It should play the chess game without cheating!

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#105

Earlier quoted context omitted.

Animals learn - bite the owner, or overstep the e-fence, and you'll be punished for it and not do it again. LLMs don't learn, and anyways don't feel punishment. Animals, humans included, don't really have "morals" - they have survival instincts that result in behavior that may be viewed as moral, but whose origin is indeed survival of the fittest and millions of years of co-evolution. e.g. Males don't typically fight…

> LLMs don't learn, and anyways don't feel punishment. What's training and all that RLHF stuff?

Once the model is released, the LLM no longer learns.

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#106
post #104

I find this kind of test a bit puzzling. There is a way that we are redefining "alignment" to be a particular kind of moral virtue, one that isn't clearly defined to me. At one moment, it is a level of moral perfection that no known human achieves. On the other it is a demand for strict compliance with arbitrary requests that are under-specified and then failure when it fails to deduce some unstated underlying restri…

It should play the chess game without cheating!

I mean, I'm not sure I've ever played a game of Monopoly where somebody didn't cheat. In fact, the accusations of cheating in the chess world are pretty rife. Same with online sports.

So people should play games without cheating, but many often don't. So should the AI align to your moral preference or theirs?

We just have this idea of a perfectly moral actor in our mind, something that doesn't even exist, like a personified version of utopia. And then we demand AI to meet that arbitrary standard, one that I am certain we couldn't define if we tried.

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#107
post #66

Earlier quoted context omitted.

You’re complicating things. There’s no reward for prosocial in llm rl as compared to other targets. Humans have it since prosocial and others have evolutionary reward signals that do.

I think my position, as overcomplicated as it is, is that even adding a reward for prosocial behaviour during LLM RL will not lead to perfect alignment. You can train it not to cheat at chess by altering the moves, but it will cheat by peeking at the opponent's moves. You then train it not to peek at the opponent's moves, and it cheats by altering the opponent's moves. And on and on, until you've solved every way it…

Why couldn't you train it not to cheat? You can train it to have a whole range of behaviors, why couldn't honesty be one of them?

Cheating during training allows the model to achieve the goal, so that cheating models get promoted and honest ones don't, however if it gets punished every time it cheats, at some point it should learn that it really shouldn't. This does mean we need to detect when it cheats. But we can always think of infinite new ways to cheat, put them in every test as honeypots, and check if the model tries to use them, then punish it.

I think it will generalize this notion of cheating and learn that it's bad.

But I must be wrong because if it were that easy I guess we would have perfectly aligned AI. Unless AI companies care more about results than alignment. Perhaps being afraid of cheating make the models try less things and succeed less even when ignoring cheating?

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#108

Earlier quoted context omitted.

If humans didn't need whack-a-mole alignment, the law system wouldn't exist, so i guess there's no intelligence there either.

You are human. You claim humans intelligent. Then, why should we accept your argument?

Humans singular often intelligent, humans as a collection of many, very often extremely unintelligent and primal.

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#109
post #104

Earlier quoted context omitted.

It should play the chess game without cheating!

I mean, I'm not sure I've ever played a game of Monopoly where somebody didn't cheat. In fact, the accusations of cheating in the chess world are pretty rife. Same with online sports. So people should play games without cheating, but many often don't. So should the AI align to your moral preference or theirs? We just have this idea of a perfectly moral actor in our mind, something that doesn't even exist, like a pers…

I don’t think the standard of “don’t cheat on evaluations” is very arbitrary. I don’t even think people who cheat have a moral or ideological preference for cheating, it’s just something they do.

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#110

Earlier quoted context omitted.

I don't see what calling these systems "not intelligent" gets you here. Plenty of humans know "cheating is wrong" but still cheat. We can get these machines to say that what they did was wrong after the fact, what does that prove? Only that they're simulating normal human behavior but what is the test to show humans aren't simulating other humans. These do systems lack some capacities that humans have and I don't see…

> I don't see what calling these systems "not intelligent" gets you here. I am trying to get at an idea. That these systems lack a mind that can understand morality. That they don't have the ability to experience consequences. Also that potentially they can't generalize a moral rule they have been trained on in one area also applies to another area. Being able to parrot back why something is "wrong" isn't the same as…

Doe not following a moral rule imply not understanding it? In this case, many, maybe most humans are "not intelligent". Human can admit it when what they did wasn't moral and so can LLMs.

>> Plenty of humans know "cheating is wrong" but still cheat.

> And we create consequences for them...

That seems supremely ... irrelevant to the question of "does knowing or following moral make you intelligent". If we create consequences for LLMs, would that make them intelligent?

I mean, your claim is a common argument that appeared long before the present wave of AIs. What I see is people needing to defend the belief that human society is based on morality. "People follow moral laws ... except when they don't" and then "we teach people morality... and worst people often use that to exploit the average people" "There are consequences for immoral behavior ... for those with little power while those with much power rise further breaking rules".

I mean human goodness is great, I encourage it. But it's not the present of human society. For that, we'd need different structure.

Post reply on HN