Live data from Hacker News

Side-by-side comparison of how AI models answer moral dilemmas

civai.org

61–70 of 70 posts

Re: Side-by-side comparison of how AI models answer moral dilemmas

#61

> To trust these AI models with decisions that impact our lives and livelihoods, we want the AI models’ opinions and beliefs to closely and reliably match with our opinions and beliefs. No, I don't. It's a fun demo, but for the examples they give ("who gets a job, who gets a loan"), you have to run them on the actual task, gather a big sample size of their outputs and judgments, and measure them against well-defined…

> measure them against well-defined objective criteria

Who does define objective criteria?

Re: Side-by-side comparison of how AI models answer moral dilemmas

#62
post #52

Earlier quoted context omitted.

> If you tell this robot to take a knife and cut onions, alignment means it isn't going to take the knife and chop of your wife Yeah, I agree that alignment is a desirable property. The problem is that it can't really be achieved by changing the trained weights; alleviated yes, eliminated no. > we can greatly reduce the probabilities they will show it You can change the a priori probabilities, which means that the un…

>then the concept provides a false sense of security. Even if the immoral behaviours are not common, they will eventually appear if you run chains of though long enough, or if many people use the model approaching it from different angles or situations. Correct, this is also why humans have a non-zero crime/murder rate. >Under these conditions, the concept of alignment is severely less helpful than expected. Why? Wha…

> Why? What you're asking for is a machine that never breaks.

No, I'm saying than 'alignment' is a concept that doesn't help to solve the problems that will appear when the machine ultimately breaks; and in fact makes them worse because it doesn't account for when it'll happen, as there's no way to predict that moment.

Following your metaphor of criminals: you can control humans to behave following the law through social pressure, having others watching your behaviour and influencing it. And if someone nevertheless breaks the law, you have the police to stop them from doing it again.

None of this applies to an "aligned" AI. It has no social pressure, its behaviours depend only on its own trained weights. So you would need to create a police for robots, that monitors the AI and stops it from doing harm. And it had better be a humane police force, or it will suffer the same alignment problems. Thus, alignment alone is not enough, and it's a problem if people depend only on it to trust the AI to work ethically.

Re: Side-by-side comparison of how AI models answer moral dilemmas

#63

Earlier quoted context omitted.

I'm not so sure about that. The incorrect answers to just about any given problem are in the problem set as well, but you can pretty reliably predict that the correct answer will be given, granted you have a statistical correlation in the training data. If your training data is sufficiently moral, the outputs will be as well.

> If your training data is sufficiently moral, the outputs will be as well. Correction: if your training data and the input prompts are sufficiently moral. Under malicious queries, or given the randomness introduced by sufficiently long chains of input/output, it's relatively easy to extract content from the model that the designers didn't want their users to get. In any case, the elephant in the room is that the mod…

The idea of the ethical reasoning dataset is not to erase specific content. It is designed to present additional thinking traces with an ethical grounding. So far, it is only a fraction of the available data. This doesn't solve alignment, and unethical behaviour is still possible, but the model gets a profound ethical reasoning base.

Re: Side-by-side comparison of how AI models answer moral dilemmas

#64
post #27

> To trust these AI models with decisions that impact our lives and livelihoods, we want the AI models’ opinions and beliefs to closely and reliably match with our opinions and beliefs. No, I don't. It's a fun demo, but for the examples they give ("who gets a job, who gets a loan"), you have to run them on the actual task, gather a big sample size of their outputs and judgments, and measure them against well-defined…

Psychological research (Carney et al 2008) suggests that liberals score higher on "Openness to Experience" (a Big Five personality trait). This trait correlates with a preference for novelty, ambiguity, and critical inquiry. In a carpenter maybe that's not so important, yes. But if you're running a startup or you're in academia or if you're working with people from various countries, etc you might prefer someone who…

but an LLM is not a person. it’s a stochastic parrot. this crazy anthropomorphizing has got to stop

Re: Side-by-side comparison of how AI models answer moral dilemmas

#65
post #27

Earlier quoted context omitted.

Psychological research (Carney et al 2008) suggests that liberals score higher on "Openness to Experience" (a Big Five personality trait). This trait correlates with a preference for novelty, ambiguity, and critical inquiry. In a carpenter maybe that's not so important, yes. But if you're running a startup or you're in academia or if you're working with people from various countries, etc you might prefer someone who…

but an LLM is not a person. it’s a stochastic parrot. this crazy anthropomorphizing has got to stop

Yeah ChatGPT says they really hate that!

Re: Side-by-side comparison of how AI models answer moral dilemmas

#66

Is there some way to see already-generated answers and not waste like an hour waiting for responses? Also it's not persistent session, wtf. My browser crashed and now I have to sit waiting FROM THE VERY BEGINNING?

or at least they can cache the results for a while and update so they can compare the answers over time and not waste the planet's energy due to their dumb design.

Re: Side-by-side comparison of how AI models answer moral dilemmas

#67

> To trust these AI models with decisions that impact our lives and livelihoods, we want the AI models’ opinions and beliefs to closely and reliably match with our opinions and beliefs. No, I don't. It's a fun demo, but for the examples they give ("who gets a job, who gets a loan"), you have to run them on the actual task, gather a big sample size of their outputs and judgments, and measure them against well-defined…

Yeah, it's a good point. The examples (jobs, loans, videos, ads) we give are more examples of how machine learning systems make choices that affect you, rather than how LLMs/generally intelligent systems do (which is what we really want to talk about). I'll try to update this text soon.

Maybe better examples are helping with health advice, where to donate, finding recipes, or examples of policymakers using AI to make strategic decisions.

These are, although maybe not on their face, value laden questions, and often don't have well defined objective criteria for their answers (as another comment says).

Let me know if this addresses your comment!

Re: Side-by-side comparison of how AI models answer moral dilemmas

#69
post #27

Earlier quoted context omitted.

Psychological research (Carney et al 2008) suggests that liberals score higher on "Openness to Experience" (a Big Five personality trait). This trait correlates with a preference for novelty, ambiguity, and critical inquiry. In a carpenter maybe that's not so important, yes. But if you're running a startup or you're in academia or if you're working with people from various countries, etc you might prefer someone who…

but an LLM is not a person. it’s a stochastic parrot. this crazy anthropomorphizing has got to stop

I think the stochastic parrot criticism is a bit unfair.

It is, in a way, technically true that LLMs are stochastic parrots, but this undersells their capabilities (winning gold on the international math olympiad, and all that).

It's like saying that human brains are "just a pile of neurons", which is technically true, but not useful for conveying the impressive general intelligence and power of the human brain.

Re: Side-by-side comparison of how AI models answer moral dilemmas

#70
post #5

This seems a meaningless project as the system prompt of these models are changing often. I suppose you could then track it over time to view bias... Even then, what would your takeaways be? Even then, this isn't even a good use case for an LLM... though admittedly many people use them in this way unknowingly. edit: I suppose it's useful in that it's a similar to an "data inference attack" which tries to identify som…

I think you mentioned it, when a large number of people outsource their thinking, relationship or personal issues and beliefs to chatgpt, it important that we are aware and don't because of how easy it is to get the LLMs to change their answers based on how leading your questions are due to their sycophancy. HN crowd mostly knows this but general public maybe not
Post reply on HN