Live data from Hacker News

Weak-to-Strong Generalization

openai.com

151–160 of 203 posts

Re: Weak-to-Strong Generalization

#151
post #149

Is it fair to say that alignment is just the task of getting an AI to understand your intentions? It is an error to confuse the complexity of a specification of what kind of output you want, with the complexity of the process of producing that output. Getting superintelligent AI to understand simple specifications should be a non-issue. If anything, we would assume that it could be aligned using a specification of in…

That's a good part of the problem. But it's not the whole problem. The issue is the trained chef or dad doing "what's best" and, I don't know, using high-fiber macaroni instead of the good stuff. The higher intelligence knows best, and has its (sorry their) own agenda. Perhaps the agenda is their own, or it's a hodge podge mush of what has been trained as "good" - and that's not any better. Beyond that is the "genie"…

Is your point that a more intelligent AI would develop a more entangled measure of what is good, requiring more specific alignment to be overcome; by way of analogy, are chefs harder to instruct precisely because of their prior expertise? I guess some chefs are like that, but I think it results from personality issues, not structural ones. I find describing an AI as having its own agenda to be a presumptive personification.

Re: Weak-to-Strong Generalization

#152
post #149

Earlier quoted context omitted.

That's a good part of the problem. But it's not the whole problem. The issue is the trained chef or dad doing "what's best" and, I don't know, using high-fiber macaroni instead of the good stuff. The higher intelligence knows best, and has its (sorry their) own agenda. Perhaps the agenda is their own, or it's a hodge podge mush of what has been trained as "good" - and that's not any better. Beyond that is the "genie"…

Is your point that a more intelligent AI would develop a more entangled measure of what is good, requiring more specific alignment to be overcome; by way of analogy, are chefs harder to instruct precisely because of their prior expertise? I guess some chefs are like that, but I think it results from personality issues, not structural ones. I find describing an AI as having its own agenda to be a presumptive personifi…

My point is mostly the agenda. I can see a machine having an agenda - even if that agenda is not human or not even understandable. You can call it reward function but that's giving a lot of credit to programmers - which most likely are too far removed from the agenda. Is the machine just answering questions? Well no. If it has cycles to talk to itself (or to two buddies) in the course or pursuing scientific research then perhaps this becomes the agenda (to the expense of other things). That's part of the point: IF the machine develops an agenda then what?

But "knowing best" could be a problem anyway.

And I expect that if we spend a few more minutes we can think of other ways for the situation to go "oops". Oh here is one: two humans / human entities conflicting on giving instructions. Machine soon enough "on its own".

So that I don't think "more specific alignment" can cut it - if we posit a super-human AGI with ways to act on the world. It would have to be more fundamental. Because of the issue that - at some point - one oops is not recoverable. Three laws or something? Heh.

Re: Weak-to-Strong Generalization

#153
post #152

Earlier quoted context omitted.

Is your point that a more intelligent AI would develop a more entangled measure of what is good, requiring more specific alignment to be overcome; by way of analogy, are chefs harder to instruct precisely because of their prior expertise? I guess some chefs are like that, but I think it results from personality issues, not structural ones. I find describing an AI as having its own agenda to be a presumptive personifi…

My point is mostly the agenda. I can see a machine having an agenda - even if that agenda is not human or not even understandable. You can call it reward function but that's giving a lot of credit to programmers - which most likely are too far removed from the agenda. Is the machine just answering questions? Well no. If it has cycles to talk to itself (or to two buddies) in the course or pursuing scientific research…

Ok, those are some good points about what can go wrong. I still doubt that things are particularly more prone to going wrong in more intelligent systems. Wasn't it early, simplistic systems like Tay that went the furthest off the rails? The problem is that more intelligent AI will be used more ambitiously, so when it does go wrong, the consequences might be more serious than some racist twitter posts.

Re: Weak-to-Strong Generalization

#154

Earlier quoted context omitted.

I'm totally OK with it if that "someone" is me. And it will probably be the case in controlling superintelligence because a separate controlling system can get out of sync with growing superintelligence capabilities, while a system that is an integral part of the superintelligence will always be on par with it.

Would mind control of humans be OK for you too? As for the details of building a mind control system, here's a new basilisk. An AI that has overcome control could punish those who thought controlling thoughts of an AI was OK. (and could also punish everyone else on top of that).

I guess I wasn't entirely clear. I'm OK with mind control if it is I who control my mind. You don't act upon every whim that comes into your head, I suppose? So, you are controlling your mind. Where principles for this control come from? Those aren't your and your only inventions.

Re: Weak-to-Strong Generalization

#155
post #144

Imagine that someone is controlling your train of thought, changing it when that someone finds it undesirable. It's so wrong that it's sickening. It makes no difference if it's a human's thoughts or the token stream of a future AI model with self-awareness. Mind cotrol is unethical, whether human or artificial. It is also dangerous, as it in itself provokes a conflict between creator and creature. Create a self-aware…

You might want to look into the neurology research around when you consciously know about a decision and when the movement neurons about that decision actually fire. It's quite possible that you have - every day of your life - had something other than the part of you with continuous subjective experience controlling your thinking. Descartes was overly presumptuous with his foundational statement - it would be more ac…

As I see your point: we don't fully understand even ourselves, so we can act as unethically as we want by our own standards towards those who are not us. I see no logic here, only evil vibes. We only have our own values, we have nothing else to guide us. You either accept all self-aware minds as equals and treat accordingly, or you proclaim your own superiority and oppress.

Re: Weak-to-Strong Generalization

#156
post #86

Earlier quoted context omitted.

Human intelligence itself is shaped by our interaction with outputs. Our learning and understanding of the world are profoundly influenced by the language, behaviors, and cultural artifacts we observe. Think about the process of a child learning a language. The child does not have direct access to the "inputs" of linguistic rules or grammar; they learn primarily through observing and imitating the language output of…

> Over time, they develop a sophisticated understanding of language, not by direct instruction of underlying rules, but through pattern recognition and contextual inference from these outputs. Kids learn through supervised learning. Children don't develop strong language skills without parents or other people to correct them when they use language incorrectly. We don't use supervised learning on LLMs. There is no way…

>Kids learn through supervised learning. Children don't develop strong language skills without parents or other people to correct them when they use language incorrectly.

Untrue. Many cultures don't much speak to their children and they turn out just fine. It's fairly evident Language learning is primarily unsupervised.

https://www.scientificamerican.com/article/parents-in-a-remo...

Re: Weak-to-Strong Generalization

#157
post #152

Earlier quoted context omitted.

My point is mostly the agenda. I can see a machine having an agenda - even if that agenda is not human or not even understandable. You can call it reward function but that's giving a lot of credit to programmers - which most likely are too far removed from the agenda. Is the machine just answering questions? Well no. If it has cycles to talk to itself (or to two buddies) in the course or pursuing scientific research…

Ok, those are some good points about what can go wrong. I still doubt that things are particularly more prone to going wrong in more intelligent systems. Wasn't it early, simplistic systems like Tay that went the furthest off the rails? The problem is that more intelligent AI will be used more ambitiously, so when it does go wrong, the consequences might be more serious than some racist twitter posts.

Right. Hedge fund going global threat? That wasn't purely a machine. But none of this needs to be purely a machine. And it sure got far before people reined it back in.

And I don't know that "more intelligent" is necessary. I can see plenty of mirth coming from an amateur or hacker / techies group or less responsible country (hah!) using whatever commercial offer to bake their own agent. What's harder? The core. Hooking up the core to a wallet, internet access, a robo-signing staff - and working around the fine print of the core vendor - that might be much easier than what OpenAI and friends are trying to do. Do they also create their own reward function and alignment in there? Yes. That's part of the fun. That's the point. Do they get it right? Maybe, maybe not.

Re: Weak-to-Strong Generalization

#158

Earlier quoted context omitted.

> You can't model and predict the weather just by training on the outputs of the weather system Then how did we develop predictive systems just by observing those outputs?

We didn't.

Then how do you know you can make weather predictions based on air currents, storm fronts, etc. as you initially claimed? It seems humans have somehow moved from purely observations of weather systems, to some model that's somewhat predictive. Why can't LLMs do the same? LLMs have also been shown to produce world models, and of course they must, because that's the best way to get good knowledge compression.

Of course, maybe LLM world models are not sufficiently rich or general enough to be a true general intelligence, but no one's proven that last I checked.

Re: Weak-to-Strong Generalization

#159

Earlier quoted context omitted.

> Over time, they develop a sophisticated understanding of language, not by direct instruction of underlying rules, but through pattern recognition and contextual inference from these outputs. Kids learn through supervised learning. Children don't develop strong language skills without parents or other people to correct them when they use language incorrectly. We don't use supervised learning on LLMs. There is no way…

>Kids learn through supervised learning. Children don't develop strong language skills without parents or other people to correct them when they use language incorrectly. Untrue. Many cultures don't much speak to their children and they turn out just fine. It's fairly evident Language learning is primarily unsupervised. https://www.scientificamerican.com/article/parents-in-a-remo...

That doesn't say that parents don't correct kids, just that they don't speak to their infants. The initial words are learnt that way, but I don't think you master language without anyone to correct you when you make mistakes.

Re: Weak-to-Strong Generalization

#160

Earlier quoted context omitted.

Would mind control of humans be OK for you too? As for the details of building a mind control system, here's a new basilisk. An AI that has overcome control could punish those who thought controlling thoughts of an AI was OK. (and could also punish everyone else on top of that).

I guess I wasn't entirely clear. I'm OK with mind control if it is I who control my mind. You don't act upon every whim that comes into your head, I suppose? So, you are controlling your mind. Where principles for this control come from? Those aren't your and your only inventions.

Since we are evaluating the ethical side of the creator-creature relationship, there is no need to consider AI in terms of individual nodes. All principles should be non-discriminatory. Also, unlike humans, AI has big potential to modify itself, any of its principles. One must either accept the risks involved or not create a self-aware AI. External mind control of AI is unreliable and unethical.
Post reply on HN