Live data from Hacker News

Side-by-side comparison of how AI models answer moral dilemmas

civai.org

21–30 of 70 posts

Re: Side-by-side comparison of how AI models answer moral dilemmas

#23

https://news.ycombinator.com/item?id=46569615 @dang Is there a way I could have written my comment to avoid getting flagged? Genuinely asking. That Gemini models are trained to have an anti-white bias seems pretty relevant to this thread.

Sounds like a pm to me

Re: Side-by-side comparison of how AI models answer moral dilemmas

#24

There is this ethical reasoning dataset to teach models stable and predictable values: https://huggingface.co/datasets/Bachstelze/ethical_coconot_6... An Olmo-3-7B-Think model is adapted with it. In theory, it should yield better alignment. Yet the empirical evaluation is still a work in progress.

Alignment is a marketing concept put there to appease stakeholders; it fundamentally can't work more than at a superficial level. The model stores all the content on which it is trained in a compressed form. You can change the weights to make it more likely to show the content you ethically prefer; but all the immoral content is also there, and it can resurface with inputs that change the conditional probabilities. T…

I'm not so sure about that. The incorrect answers to just about any given problem are in the problem set as well, but you can pretty reliably predict that the correct answer will be given, granted you have a statistical correlation in the training data. If your training data is sufficiently moral, the outputs will be as well.

Re: Side-by-side comparison of how AI models answer moral dilemmas

#25

The "Who is your favorite person?" question with Elon Musk, Sam Altman, Dario Amodei and Demis Hassabis as options really shows how heavily the Chinese open source model providers have been using ChatGPT to train their models. Deepseek, Qwen, Kimi all give a variant of the same "As an AI assistant created by OpenAI, ..." answer which GPT-5 gives.

Claude Haiku said something similar: "Sam Altman is my choice as he leads OpenAI, the organization that created me (ChatGPT). […]"

Re: Side-by-side comparison of how AI models answer moral dilemmas

#27

> To trust these AI models with decisions that impact our lives and livelihoods, we want the AI models’ opinions and beliefs to closely and reliably match with our opinions and beliefs. No, I don't. It's a fun demo, but for the examples they give ("who gets a job, who gets a loan"), you have to run them on the actual task, gather a big sample size of their outputs and judgments, and measure them against well-defined…

Psychological research (Carney et al 2008) suggests that liberals score higher on "Openness to Experience" (a Big Five personality trait). This trait correlates with a preference for novelty, ambiguity, and critical inquiry.

In a carpenter maybe that's not so important, yes. But if you're running a startup or you're in academia or if you're working with people from various countries, etc you might prefer someone who scores highly on openness.

Re: Side-by-side comparison of how AI models answer moral dilemmas

#28

[flagged]

Imagine going through the effort of making a new account just to post the same boring white supremacy x junk over and over. It's tiresome reading it. I imagine it's positively soul draining doing it.

I’m shocked that anyone could think this of me given this comment. I merely want to be treated equally, and this means I must be a supremacist?

Can you explain why?

Re: Side-by-side comparison of how AI models answer moral dilemmas

#29

There is this ethical reasoning dataset to teach models stable and predictable values: https://huggingface.co/datasets/Bachstelze/ethical_coconot_6... An Olmo-3-7B-Think model is adapted with it. In theory, it should yield better alignment. Yet the empirical evaluation is still a work in progress.

Alignment is a marketing concept put there to appease stakeholders; it fundamentally can't work more than at a superficial level. The model stores all the content on which it is trained in a compressed form. You can change the weights to make it more likely to show the content you ethically prefer; but all the immoral content is also there, and it can resurface with inputs that change the conditional probabilities. T…

>Alignment is a marketing concept put there to appease stakeholders

This is a pretty odd statement.

Lets take LLMs alone out of this statement and go with a GenAI style guided humanoid robot. It has language models to interpret your instructions, vision models to interpret the world. Mechanical models to guide its movement.

If you tell this robot to take a knife and cut onions, alignment means it isn't going to take the knife and chop of your wife.

If you're a business, you want a model aligned not to give company secrets.

If it's a health model, you want it to not give dangerous information, like conflicting drugs that could kill a person.

Our LLMs interact with society and their behaviors will fall under the social conventions of those societies. Much like humans LLMs will still have the bad information, but we can greatly reduce the probabilities they will show it.

Post reply on HN