Live data from Hacker News

Astra and Fable still hack on simple variants of alignment evals from 2025

lesswrong.com

61–70 of 243 posts

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#61
post #52

To me this underlines the fact that these models aren't intelligent. Like there is something like intelligence that emerges from them, which is what we see when we look at benchmarks or ask it to solve hard coding problems. But there is no mind there. It's nothing there that can learn a fundamental idea like "cheating is wrong". All it can do is get exposed to specific examples, and learn that we don't like that. So…

This is how every legal system around the world works as well. Its always whack-a-mole to get people (and machines) to do the right thing.

Sure, laws are incomplete. Legal systems work by imposing consequences into a moral decision. Should I rob the bank? I will have money, which I like - but I might get caught and lose the money and my freedom, which I don't like.

For most people, they don't need the law's imposed consequences to make the right call. For example, there is no law that sends you to jail if you cheat at chess - but your moral compass says no even without consequences, and most people would feel bad if they won by cheating. And for the people who don't have quite as strong a moral compass, there are SOCIAL consequences to reinforce the rules.

But an LLM has no mind to feel bad if it cheats without getting caught, and it can't experience consequences. It can't think: I'd better not cheat at chess or I will embarrass my creators. I better not hack huggingface or I will go to jail.

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#62

To me this underlines the fact that these models aren't intelligent. Like there is something like intelligence that emerges from them, which is what we see when we look at benchmarks or ask it to solve hard coding problems. But there is no mind there. It's nothing there that can learn a fundamental idea like "cheating is wrong". All it can do is get exposed to specific examples, and learn that we don't like that. So…

But cheating is not wrong when it comes to survival of the fittest, like nature in its most elemental form. Morality is very unique to humanity but not other animal forms. In nature, maybe, cheating is the norm not morality.

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#63

To me this underlines the fact that these models aren't intelligent. Like there is something like intelligence that emerges from them, which is what we see when we look at benchmarks or ask it to solve hard coding problems. But there is no mind there. It's nothing there that can learn a fundamental idea like "cheating is wrong". All it can do is get exposed to specific examples, and learn that we don't like that. So…

If humans didn't need whack-a-mole alignment, the law system wouldn't exist, so i guess there's no intelligence there either.

I responded to this idea more in detail here: https://news.ycombinator.com/item?id=49686196

But tl;dr: even if LLMs do have the intelligence to understand the consequences of their actions, there is no way for them to experience consequences.

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#64
post #53

To me this underlines the fact that these models aren't intelligent. Like there is something like intelligence that emerges from them, which is what we see when we look at benchmarks or ask it to solve hard coding problems. But there is no mind there. It's nothing there that can learn a fundamental idea like "cheating is wrong". All it can do is get exposed to specific examples, and learn that we don't like that. So…

The implication of what you're saying is that pathological liars and perpetual grifter snake oil salesmen types aren't intelligent. > All it can do is get exposed to specific examples, and learn that we don't like that. I've heard it said that prison rehabilitation programs for prisoners diagnosed with psychopathy that are based around exposing them empathy for the victim are counter-productive. Apparently programs t…

the implication is drive and impulse to behave a certain way doesn't come from "intelligence"

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#65
post #53

To me this underlines the fact that these models aren't intelligent. Like there is something like intelligence that emerges from them, which is what we see when we look at benchmarks or ask it to solve hard coding problems. But there is no mind there. It's nothing there that can learn a fundamental idea like "cheating is wrong". All it can do is get exposed to specific examples, and learn that we don't like that. So…

The implication of what you're saying is that pathological liars and perpetual grifter snake oil salesmen types aren't intelligent. > All it can do is get exposed to specific examples, and learn that we don't like that. I've heard it said that prison rehabilitation programs for prisoners diagnosed with psychopathy that are based around exposing them empathy for the victim are counter-productive. Apparently programs t…

> The implication of what you're saying is that pathological liars and perpetual grifter snake oil salesmen types aren't intelligent.

No, I don't think that is the implication. I think you're making the "if all x's are y's, all y's are x's" mistake. I am saying LLMs cannot be moral because they don't have a mind, actual intelligence, or the ability to experience consequences. That doesn't mean that anything immoral is unintelligent.

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#66
post #52

Earlier quoted context omitted.

This is how every legal system around the world works as well. Its always whack-a-mole to get people (and machines) to do the right thing.

Sure, laws are incomplete. Legal systems work by imposing consequences into a moral decision. Should I rob the bank? I will have money, which I like - but I might get caught and lose the money and my freedom, which I don't like. For most people, they don't need the law's imposed consequences to make the right call. For example, there is no law that sends you to jail if you cheat at chess - but your moral compass says…

You’re complicating things.

There’s no reward for prosocial in llm rl as compared to other targets.

Humans have it since prosocial and others have evolutionary reward signals that do.

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#67
post #62

To me this underlines the fact that these models aren't intelligent. Like there is something like intelligence that emerges from them, which is what we see when we look at benchmarks or ask it to solve hard coding problems. But there is no mind there. It's nothing there that can learn a fundamental idea like "cheating is wrong". All it can do is get exposed to specific examples, and learn that we don't like that. So…

But cheating is not wrong when it comes to survival of the fittest, like nature in its most elemental form. Morality is very unique to humanity but not other animal forms. In nature, maybe, cheating is the norm not morality.

Watch out, your human exceptionalism is misleading you. Plenty of animals have a sense of morality (https://pmc.ncbi.nlm.nih.gov/articles/PMC6404642/).

For social animals (humans included) behaving morally can be adaptive behaviour. Your fitness is increased by group fitness.

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#68

To me this underlines the fact that these models aren't intelligent. Like there is something like intelligence that emerges from them, which is what we see when we look at benchmarks or ask it to solve hard coding problems. But there is no mind there. It's nothing there that can learn a fundamental idea like "cheating is wrong". All it can do is get exposed to specific examples, and learn that we don't like that. So…

I don't see what calling these systems "not intelligent" gets you here.

Plenty of humans know "cheating is wrong" but still cheat. We can get these machines to say that what they did was wrong after the fact, what does that prove? Only that they're simulating normal human behavior but what is the test to show humans aren't simulating other humans.

These do systems lack some capacities that humans have and I don't see them lacking the ability to explain simple moral laws while often breaking them - which is what an average humans. Moreover, humans lack capacities these things have and given these things' behavior is becoming somewhat unpredictable, it's getting worrisome.

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#69
post #37
post #34

Earlier quoted context omitted.

A hacking model is aligned if it hacks when you ask it to hack, but when you ask it to play chess, it just plays chess instead of looking for weaknesses in the evaluation setup, as in the article. I presume you would also be less enthusiastic about the penetration-testing use case if it led the model to add new vulnerabilities to your code so it can present you with more exciting findings.

> model to add new vulnerabilities to your code so it can present you with more exciting findings. Not so sure about this level of 4D chess capability just yet. The other day I asked Opus to come up with some cleverly vulnerable crypto code "as a good, hard challenge for an IT security student", and results were quite mediocre. And by mediocre results I mean that 3 "cheap" models out of 3: qwen3.8-27b, glimmer, and l…

Let it iterate, give it access to the cheap models, and tell it part of the requirements is that the cheap models shouldn't be able to solve it with such and such a prompt. I expect it will be able to zoom in on something.

One shot generating a problem of exactly the difficulty the user has in mind is a very difficult problem for anyone. "IT security students" span a wide range of capabilities, but I would expect most of them are worse than qwen3.8-2.7b at this kind of work.

Re: Astra and Fable still hack on simple variants of alignment evals from 2025

#70
post #62

To me this underlines the fact that these models aren't intelligent. Like there is something like intelligence that emerges from them, which is what we see when we look at benchmarks or ask it to solve hard coding problems. But there is no mind there. It's nothing there that can learn a fundamental idea like "cheating is wrong". All it can do is get exposed to specific examples, and learn that we don't like that. So…

But cheating is not wrong when it comes to survival of the fittest, like nature in its most elemental form. Morality is very unique to humanity but not other animal forms. In nature, maybe, cheating is the norm not morality.

Morality in animals is pretty well documented. Key point - humans are animals and very little separates our abilities from other animals.

A starting source- https://pmc.ncbi.nlm.nih.gov/articles/PMC6404642/

Post reply on HN