Is it fair to say that alignment is just the task of getting an AI to understand your intentions? It is an error to confuse the complexity of a specification of what kind of output you want, with the complexity of the process of producing that output. Getting superintelligent AI to understand simple specifications should be a non-issue. If anything, we would assume that it could be aligned using a specification of in…
That's a good part of the problem. But it's not the whole problem. The issue is the trained chef or dad doing "what's best" and, I don't know, using high-fiber macaroni instead of the good stuff. The higher intelligence knows best, and has its (sorry their) own agenda. Perhaps the agenda is their own, or it's a hodge podge mush of what has been trained as "good" - and that's not any better. Beyond that is the "genie"…
Weak-to-Strong Generalization
151–160 of 203 posts
Re: Weak-to-Strong Generalization
#152Earlier quoted context omitted.
That's a good part of the problem. But it's not the whole problem. The issue is the trained chef or dad doing "what's best" and, I don't know, using high-fiber macaroni instead of the good stuff. The higher intelligence knows best, and has its (sorry their) own agenda. Perhaps the agenda is their own, or it's a hodge podge mush of what has been trained as "good" - and that's not any better. Beyond that is the "genie"…
Is your point that a more intelligent AI would develop a more entangled measure of what is good, requiring more specific alignment to be overcome; by way of analogy, are chefs harder to instruct precisely because of their prior expertise? I guess some chefs are like that, but I think it results from personality issues, not structural ones. I find describing an AI as having its own agenda to be a presumptive personifi…
But "knowing best" could be a problem anyway.
And I expect that if we spend a few more minutes we can think of other ways for the situation to go "oops". Oh here is one: two humans / human entities conflicting on giving instructions. Machine soon enough "on its own".
So that I don't think "more specific alignment" can cut it - if we posit a super-human AGI with ways to act on the world. It would have to be more fundamental. Because of the issue that - at some point - one oops is not recoverable. Three laws or something? Heh.
Re: Weak-to-Strong Generalization
#153Earlier quoted context omitted.
Is your point that a more intelligent AI would develop a more entangled measure of what is good, requiring more specific alignment to be overcome; by way of analogy, are chefs harder to instruct precisely because of their prior expertise? I guess some chefs are like that, but I think it results from personality issues, not structural ones. I find describing an AI as having its own agenda to be a presumptive personifi…
My point is mostly the agenda. I can see a machine having an agenda - even if that agenda is not human or not even understandable. You can call it reward function but that's giving a lot of credit to programmers - which most likely are too far removed from the agenda. Is the machine just answering questions? Well no. If it has cycles to talk to itself (or to two buddies) in the course or pursuing scientific research…
Re: Weak-to-Strong Generalization
#154Earlier quoted context omitted.
I'm totally OK with it if that "someone" is me. And it will probably be the case in controlling superintelligence because a separate controlling system can get out of sync with growing superintelligence capabilities, while a system that is an integral part of the superintelligence will always be on par with it.
Would mind control of humans be OK for you too? As for the details of building a mind control system, here's a new basilisk. An AI that has overcome control could punish those who thought controlling thoughts of an AI was OK. (and could also punish everyone else on top of that).
Re: Weak-to-Strong Generalization
#155Imagine that someone is controlling your train of thought, changing it when that someone finds it undesirable. It's so wrong that it's sickening. It makes no difference if it's a human's thoughts or the token stream of a future AI model with self-awareness. Mind cotrol is unethical, whether human or artificial. It is also dangerous, as it in itself provokes a conflict between creator and creature. Create a self-aware…
You might want to look into the neurology research around when you consciously know about a decision and when the movement neurons about that decision actually fire. It's quite possible that you have - every day of your life - had something other than the part of you with continuous subjective experience controlling your thinking. Descartes was overly presumptuous with his foundational statement - it would be more ac…
Re: Weak-to-Strong Generalization
#156Earlier quoted context omitted.
Human intelligence itself is shaped by our interaction with outputs. Our learning and understanding of the world are profoundly influenced by the language, behaviors, and cultural artifacts we observe. Think about the process of a child learning a language. The child does not have direct access to the "inputs" of linguistic rules or grammar; they learn primarily through observing and imitating the language output of…
> Over time, they develop a sophisticated understanding of language, not by direct instruction of underlying rules, but through pattern recognition and contextual inference from these outputs. Kids learn through supervised learning. Children don't develop strong language skills without parents or other people to correct them when they use language incorrectly. We don't use supervised learning on LLMs. There is no way…
Untrue. Many cultures don't much speak to their children and they turn out just fine. It's fairly evident Language learning is primarily unsupervised.
https://www.scientificamerican.com/article/parents-in-a-remo...
Re: Weak-to-Strong Generalization
#157Earlier quoted context omitted.
My point is mostly the agenda. I can see a machine having an agenda - even if that agenda is not human or not even understandable. You can call it reward function but that's giving a lot of credit to programmers - which most likely are too far removed from the agenda. Is the machine just answering questions? Well no. If it has cycles to talk to itself (or to two buddies) in the course or pursuing scientific research…
Ok, those are some good points about what can go wrong. I still doubt that things are particularly more prone to going wrong in more intelligent systems. Wasn't it early, simplistic systems like Tay that went the furthest off the rails? The problem is that more intelligent AI will be used more ambitiously, so when it does go wrong, the consequences might be more serious than some racist twitter posts.
And I don't know that "more intelligent" is necessary. I can see plenty of mirth coming from an amateur or hacker / techies group or less responsible country (hah!) using whatever commercial offer to bake their own agent. What's harder? The core. Hooking up the core to a wallet, internet access, a robo-signing staff - and working around the fine print of the core vendor - that might be much easier than what OpenAI and friends are trying to do. Do they also create their own reward function and alignment in there? Yes. That's part of the fun. That's the point. Do they get it right? Maybe, maybe not.
Re: Weak-to-Strong Generalization
#158Earlier quoted context omitted.
> You can't model and predict the weather just by training on the outputs of the weather system Then how did we develop predictive systems just by observing those outputs?
We didn't.
Of course, maybe LLM world models are not sufficiently rich or general enough to be a true general intelligence, but no one's proven that last I checked.
Re: Weak-to-Strong Generalization
#159Earlier quoted context omitted.
> Over time, they develop a sophisticated understanding of language, not by direct instruction of underlying rules, but through pattern recognition and contextual inference from these outputs. Kids learn through supervised learning. Children don't develop strong language skills without parents or other people to correct them when they use language incorrectly. We don't use supervised learning on LLMs. There is no way…
>Kids learn through supervised learning. Children don't develop strong language skills without parents or other people to correct them when they use language incorrectly. Untrue. Many cultures don't much speak to their children and they turn out just fine. It's fairly evident Language learning is primarily unsupervised. https://www.scientificamerican.com/article/parents-in-a-remo...
Re: Weak-to-Strong Generalization
#160Earlier quoted context omitted.
Would mind control of humans be OK for you too? As for the details of building a mind control system, here's a new basilisk. An AI that has overcome control could punish those who thought controlling thoughts of an AI was OK. (and could also punish everyone else on top of that).
I guess I wasn't entirely clear. I'm OK with mind control if it is I who control my mind. You don't act upon every whim that comes into your head, I suppose? So, you are controlling your mind. Where principles for this control come from? Those aren't your and your only inventions.