Live data from Hacker News

Adding guardrails to large language models

github.com

11–15 of 15 posts

Re: Adding guardrails to large language models

#12

Wouldn’t the right way to create AI guardrails is to to have an antagonistic AI act as a moderator? Like you have one model trained to be as accurate as possible in fulfilling the prompt, and then another AI trained based on how human moderators apply the terms of another, “moderation” prompt. Then you have the two fight on a large training set and when you’re done you have generated a moderated AI.

Something like a GAN?

Will probably end up something similar though.

Re: Adding guardrails to large language models

#13

Wouldn’t the right way to create AI guardrails is to to have an antagonistic AI act as a moderator? Like you have one model trained to be as accurate as possible in fulfilling the prompt, and then another AI trained based on how human moderators apply the terms of another, “moderation” prompt. Then you have the two fight on a large training set and when you’re done you have generated a moderated AI.

To be effective the moderator AI would need to be as smart (or smarter) as the source AI. Think of all the ways we have already seen people get around restrictions. Giving instructions for murder isn't allowed but people said they were writing a novel and want to have a murder in it and how could be be done. A smart moderator would see what the user is trying to do and stop it.
Post reply on HN