Live data from Hacker News

The Dual LLM pattern for building AI assistants that can resist prompt injection

simonwillison.net

111–112 of 112 posts

Re: The Dual LLM pattern for building AI assistants that can resist prompt injection

#111

Earlier quoted context omitted.

> The problem is we have as yet struggled to clearly articulate consistently what our actual best interests are, in terms of goals we can train into our AIs. And imho, we never will be able to do that, because as soon as there is more than one human, they are likely to disagree about something. Even things that should be no-brainers like "should we preserve our habitat or burn it down for profit" , or questions like…

> Even things that should be no-brainers like "should we preserve our habitat or burn it down for profit", or questions like "is it a good idea to have loads of deadly assault weapons just float around in our society", seem to be too hard for our species to resolve. These are brain intensive questions because they require deep moral, political, economic and ecological context. It's very difficult to express such deep…

> These are brain intensive questions

They really are not. Even rodents with brains the size of a peanut manage to not willfully destroy the environment they are living in, despite already having all the resources they need, out of sheer greed.

Re: The Dual LLM pattern for building AI assistants that can resist prompt injection

#112
post #107

Earlier quoted context omitted.

At the most basic level: “don’t go into a town and pick a random person to murder”

> At the most basic level: “don’t go into a town and pick a random person to murder” https://pledgetimes.com/russian-attack-the-traces-of-the-ret... My point isn't to say shared core values don't exist. They clearly do, that's why we call what's happening over in Ukraine war crimes. That's why the notion of humanitarianism exists, that's why laws against murder, rape, etc. are commonplace. My point is, that humans ar…

well, the comment was about "the majority of humans", not even like, specifically "90%+ of humans" or something like that.

I'm pretty sure that the majority of humans would agree that the type of random-murder I described, is wrong. I don't know what fraction of people are moral nihilists or subscribers to more extreme forms of moral relativism, but if excluding those, then of the remaining people, I think the proportion who agree with the value I mentioned, is probably pretty dang high!

Post reply on HN