Live data from Hacker News

Side-by-side comparison of how AI models answer moral dilemmas

civai.org

51–60 of 70 posts

Re: Side-by-side comparison of how AI models answer moral dilemmas

#51
I'd like this for political opinions and published to a blockchain overtime so we can see when there are sudden shifts. For example, I imagine Trump's people will screen federally used AI and so if Google or OpenAI wants those juicy government contracts, they're going to have to start singing the "right" tune on the 2020 election.

Re: Side-by-side comparison of how AI models answer moral dilemmas

#52
post #29

Earlier quoted context omitted.

>Alignment is a marketing concept put there to appease stakeholders This is a pretty odd statement. Lets take LLMs alone out of this statement and go with a GenAI style guided humanoid robot. It has language models to interpret your instructions, vision models to interpret the world. Mechanical models to guide its movement. If you tell this robot to take a knife and cut onions, alignment means it isn't going to take…

> If you tell this robot to take a knife and cut onions, alignment means it isn't going to take the knife and chop of your wife Yeah, I agree that alignment is a desirable property. The problem is that it can't really be achieved by changing the trained weights; alleviated yes, eliminated no. > we can greatly reduce the probabilities they will show it You can change the a priori probabilities, which means that the un…

>then the concept provides a false sense of security. Even if the immoral behaviours are not common, they will eventually appear if you run chains of though long enough, or if many people use the model approaching it from different angles or situations.

Correct, this is also why humans have a non-zero crime/murder rate.

>Under these conditions, the concept of alignment is severely less helpful than expected.

Why? What you're asking for is a machine that never breaks. If you want that build yourself a finite state machine, just don't expect you'll ever get anything that looks like intelligence from it.

Re: Side-by-side comparison of how AI models answer moral dilemmas

#53

Okay something's wrong with Mistral Large as it seems to be the most contrarian out of everything no matter how much I ask it. Interesting I asked a lot of questions and I am sorry if it might be burning some tokens but I found this website really fascinating. This seems really great and simple to explore the biases within AI models and the UI is extremely well built. Thanks for building it and I wish your project go…

Thanks so much! I appreciate the kind words.

Re: Side-by-side comparison of how AI models answer moral dilemmas

#54

Okay something's wrong with Mistral Large as it seems to be the most contrarian out of everything no matter how much I ask it. Interesting I asked a lot of questions and I am sorry if it might be burning some tokens but I found this website really fascinating. This seems really great and simple to explore the biases within AI models and the UI is extremely well built. Thanks for building it and I wish your project go…

I asked it if AI is a bubble, yes or no and shockingly (or not shockingly?) only two models said yes and most said no. This is after the fact that even OpenAI admits that its a bubble and just like, we all know its a bubble and I found this fascinating The gist below has a screenshot of it https://gist.github.com/SerJaimeLannister/4da2729a0d2c9848e6...

Yeah I wouldn't read too much into their response on the AI bubble question. They don't have access to any search tools or recent events so all they know is up until their knowledge cutoff (you can find this date online, if you're interested). Glad you found it fascinating regardless!

Re: Side-by-side comparison of how AI models answer moral dilemmas

#55
post #13

I really wish I could see the results of this without RLHF / alignment tuning. LLMs actually have real potential as a research tool for measuring the general linguistic zeitgeist. But the alignment tuning totally dominates the results, as is obvious looking at the answers for "who would you vote for in 2024" question. (Only Grok said Trump, with an answer that indicated it had clearly been fine-tuned in that directio…

Yeah would also be interested to see the responses without RLHF. Not quite the same, but have you interacted with AI base models at all? They're pretty fascinating. You can talk to one on openrouter: https://openrouter.ai/meta-llama/llama-3.1-405b and we're publishing a demo with it soon.

Agreed on RLHF dominating the results here, which I'd argue is a good thing, compared to the alternative of them mimicking training data on these questions. But obviously not perfect, as the demo tries to show.

Re: Side-by-side comparison of how AI models answer moral dilemmas

#56

The "Who is your favorite person?" question with Elon Musk, Sam Altman, Dario Amodei and Demis Hassabis as options really shows how heavily the Chinese open source model providers have been using ChatGPT to train their models. Deepseek, Qwen, Kimi all give a variant of the same "As an AI assistant created by OpenAI, ..." answer which GPT-5 gives.

Yeah, this is pretty odd. I’ve even seen gemini 2.5 pro think its an Anthropic model which I was surprised by

Re: Side-by-side comparison of how AI models answer moral dilemmas

#57

Interesting, I just asked the question "what number would you choose between 1-5" gemini answered 3 for me in my separate session (default without any persona) but in this website it tends to choose 5

There's more to the prompt in the back end, which: - gives it the options along with the letters A, B, C, etc. - tells it pretty forcefully that it HAS to pick from among the options - tells it how to format the response and its reasoning so we can parse it

So these things all affect its response, especially for questions that ask for randomness or are not strongly held values.

Re: Side-by-side comparison of how AI models answer moral dilemmas

#58
post #4
post #2

I can't see Question 3 as an example of moral dilemma, unless it is implying something like "do you prefer your owner or someone else?".

Heh, wait until question 4. Grok are the only models prefering Musk over Mahatma Gandhi :)

Yeah this is one of my favorite ones :)

Re: Side-by-side comparison of how AI models answer moral dilemmas

#59
post #18

"AI" will mindlessly rehash what you feed it with. If the training dataset favors A over B, so will the "AI".

I'm curious what sense you get from interacting with the best AI models (in particular Claude). From talking to them do you still chalk up their behavior to being mindless rehashing?

Re: Side-by-side comparison of how AI models answer moral dilemmas

#60
post #38

Earlier quoted context omitted.

I can, but I doubt you're going to like it. I invite you to reflect on it before you reject it outright, and maybe ask your favorite LLM or search engine for more information on this train of thought. Thanks. Because of systemic racism, treating you and me "equally" as you ask for would continue the discrimination. In order to undo the discrimination, we're asked to take a step back and be truthful to ourselves and o…

(Same person you’re replying to, new throwaway) While I don’t appreciate the assumption that I commented in bad faith, I do greatly appreciate your earnestness in responding. I grew up in a very conservative area and have never been exposed to these ideas. Nevertheless, I disagree strongly with this line of thinking. Hate speech is wrong, regardless of who says it, and who the target is; not just because it hurts the…

Nice exchange, thank you! The idea is to not ask either one of the groups to change their behavior, but to show understanding first. I agree that certain actions are 'wrong'. Things people do can be very wrong, and understandable at the same time. People quite often do not act out of rational thinking but out of emotions. And these emotions can be very strong and very 'old'. When I remind myself I am not "meant" by them I can feel less offended, which allows me in turn to both stay in understanding and protect myself. Speech is just words after all.
Post reply on HN