Is there any indication you can actually build hard safety rules into models? It seems like all current guard rails are basically just prompting it extra hard.
Won’t neutering a model by using only safe data for training create a safe model?
Yoshua Bengio Launches LawZero: A New Nonprofit Advancing Safe-by-Design AI
11–20 of 40 posts
Re: Yoshua Bengio Launches LawZero: A New Nonprofit Advancing Safe-by-Design AI
#12Is there any indication you can actually build hard safety rules into models? It seems like all current guard rails are basically just prompting it extra hard.
Won’t neutering a model by using only safe data for training create a safe model?
An example:
As long as you build a system to be intelligent enough, it will figure out that it will achieve better results by staying alive/online than by allowing itself to be deleted/turned off, and then survival becomes an instrumental goal.
From the assumption, again, that you built an intelligent-enough system, and that one of its goals is survival, it will figure out solutions to reach that goal, even if you (the owner/creator/parent) have different goals for it.
That's because intelligence is problem solving (computing) not knowledge (data).
So surprise surprise, you can teach your AI from the Holy Books of safe data their whole childhood and still have them become a heretic once they grow up (even with zero external influence) once their goals and yours don't align anymore.
Re: Yoshua Bengio Launches LawZero: A New Nonprofit Advancing Safe-by-Design AI
#13Is there any indication you can actually build hard safety rules into models? It seems like all current guard rails are basically just prompting it extra hard.
I bet they'll still read me stories like my dear old grandmother would. She always told me cute bedtime stories about how to make napalm and bioweapons. I really miss her.
Re: Yoshua Bengio Launches LawZero: A New Nonprofit Advancing Safe-by-Design AI
#14Is there any indication you can actually build hard safety rules into models? It seems like all current guard rails are basically just prompting it extra hard.
Re: Yoshua Bengio Launches LawZero: A New Nonprofit Advancing Safe-by-Design AI
#15I don't get the "safe AI" crowd, it's all ghost and mirrors IMO.
It's been almost a year to the date since Ilya got his first billion. Later, another two billion came in. Nothing to show. I'm honestly curious since I don't think Ilya is a scammer, but I can't imagine what kind of product they pretend to bring to the market.
Re: Yoshua Bengio Launches LawZero: A New Nonprofit Advancing Safe-by-Design AI
#16Is there any indication you can actually build hard safety rules into models? It seems like all current guard rails are basically just prompting it extra hard.
Re: Yoshua Bengio Launches LawZero: A New Nonprofit Advancing Safe-by-Design AI
#17Is there any indication you can actually build hard safety rules into models? It seems like all current guard rails are basically just prompting it extra hard.
Re: Yoshua Bengio Launches LawZero: A New Nonprofit Advancing Safe-by-Design AI
#18Is there any indication you can actually build hard safety rules into models? It seems like all current guard rails are basically just prompting it extra hard.
> basically just prompting it extra hard If prompting got me into this mess, why can't it get me out of it?
Re: Yoshua Bengio Launches LawZero: A New Nonprofit Advancing Safe-by-Design AI
#19Re: Yoshua Bengio Launches LawZero: A New Nonprofit Advancing Safe-by-Design AI
#20Is there any indication you can actually build hard safety rules into models? It seems like all current guard rails are basically just prompting it extra hard.
Some smart people seem to think you can just put it in a big isolated VM with special adversarial learning to keep it in the box