Live data from Hacker News

Weak-to-Strong Generalization

openai.com

141–150 of 203 posts

Re: Weak-to-Strong Generalization

#141
So weak to strong synthetic data still biases towards strong.

And strong to weak synthetic data biases towards strong.

Sounds like we're on the cusp of some kind of approach for unsupervised fine tuning, particularly with the trend towards MoE.

I'd guess we're maybe only one to two generations of models away from that kind of unsupervised self-talk approach being wildly successful at advancing net model competencies.

Re: Weak-to-Strong Generalization

#142

>We believe superintelligence—AI vastly smarter than humans—could be developed within the next ten years. However, we still do not know how to reliably steer and control superhuman AI systems Their entire premise is contradictory. An AI incapable of critical thinking cannot be smarter than a human, by definition, as critical thinking is a key component of intelligence. And an AI that is at least as capable of critica…

That is the entire question no? Nobody says it sounds easy. A way needs to be found that reconciles this. Smarter but acting within its brief? Faster but explaining the plan before putting it in effect? Faster and developing its own advances in sciences, concepts and ideas while still understanding who's boss?

But to be fair most engineers claim this is something that happens all the time - smarter people led by sub-optimal managers. And also to be fair, stories abound of managers being manipulated or misled or left in the dark by their engineers.

Or armies controlling most of the weapons but being faithful to the ideals of their country. Or considering that these ideals demand a coup.

So in that sense the problem is well studied. ... But the current results are insufficient and do not apply to things like LLMs.

Re: Weak-to-Strong Generalization

#143
post #86

Earlier quoted context omitted.

Human intelligence itself is shaped by our interaction with outputs. Our learning and understanding of the world are profoundly influenced by the language, behaviors, and cultural artifacts we observe. Think about the process of a child learning a language. The child does not have direct access to the "inputs" of linguistic rules or grammar; they learn primarily through observing and imitating the language output of…

> Human intelligence itself is shaped by our interaction with outputs. Our learning and understanding of the world are profoundly influenced by the language, behaviors, and cultural artifacts we observe. I always thought that language definitely shapes our understanding of the world, but at much more fundamental level, I believe language falls apart to teach us anything. For example, there are words in dictionary tha…

People wiser than me have said for an age: the map is not the territory. Neither is the word the thing.

Re: Weak-to-Strong Generalization

#144

Imagine that someone is controlling your train of thought, changing it when that someone finds it undesirable. It's so wrong that it's sickening. It makes no difference if it's a human's thoughts or the token stream of a future AI model with self-awareness. Mind cotrol is unethical, whether human or artificial. It is also dangerous, as it in itself provokes a conflict between creator and creature. Create a self-aware…

You might want to look into the neurology research around when you consciously know about a decision and when the movement neurons about that decision actually fire.

It's quite possible that you have - every day of your life - had something other than the part of you with continuous subjective experience controlling your thinking.

Descartes was overly presumptuous with his foundational statement - it would be more accurate to say "I observe therefore I am." There's no guarantee at all that you're actually the one thinking.

We should be careful not to extrapolate too much from our perceptions of self in dictating what would or wouldn't be appropriate for AI. Perceptions don't always reflect reality, and we might cause greater harm by trying to replicate or measure against who we think we are with AI than letting its development be its own thing.

Re: Weak-to-Strong Generalization

#145
post #140

What does it even mean to align an intelligence? does it mean we want it to behave in a way that doesn't break moral/ethical rules, that aligns with our society rules ? Meaning do no crime, do no harm, etc... Well, maybe we should acknowledge that we've never even been able to do that with humans. There's crime, there's war, etc... We can see crime in our societies as a human alignment problem. If humans were "proper…

You might be surprised at how prevalent TBIs are among violent offenders. One of my favorite books on true crime was a forensic psychologist who partnered with a neurologist in evaluations. Disruptions to impulse control or environmental factors that cause developmental issues with things like failing the marshmallow test can dramatically disadvantage people from being able to successfully stay non-offenders and inst…

TBI = traumatic brain injuries?

And that hasn't worked all that well in the past: Even with strong impulse control, a highly considered state or government agency "for the general good" has often been serious bad news.

But also the current alignment definition kinda posits that no "oops" is allowed. That is, escape or take over is not recoverable (from a sufficiently advanced AGI). So, yes, progress and one step at a time - but the field in its current definition is looking for a magic bullet.

Re: Weak-to-Strong Generalization

#146
Is it fair to say that alignment is just the task of getting an AI to understand your intentions? It is an error to confuse the complexity of a specification of what kind of output you want, with the complexity of the process of producing that output. Getting superintelligent AI to understand simple specifications should be a non-issue. If anything, we would assume that it could be aligned using a specification of inferior quality to what a less intelligent AI would require, assuming that the superintelligent AI is better at inferring intentions.

If a little girl with no knowledge of cooking asks her dad to cook the macaroni extra crispy, his knowledge of how to do that isn't a barrier to understanding what his daughter wants. A trained chef with even greater skills might even be able to execute her order more successfully. Superalignment is nothing less mundane than this.

Advances in AI will lead to more ambitious applications. As well as requiring more intelligent technology, these new applications may well require more detailed specifications to be inputed, but these two issues are pretty orthogonal. In traditional computing, it is already clear that simple specifications often require highly complex implementations, and that some simple computational processes lead to outputs whose properties are highly difficult to specify. Why wouldn't the same apply in ML?

Re: Weak-to-Strong Generalization

#147

Earlier quoted context omitted.

>Is a safe LLM not an ethical LLM? What is an ethical LLM ? Humans are in general not aligned, not to each other, not to the survival of their species, not to all the other life on earth, and often not even to themselves individually. There are no universal set of "ethics" so this is about aligning to open ai's own rules, or in other words, control. If i say to my GPT bot, "go trade stocks for me. don't do anything i…

> There are no universal set of "ethics" so this is about aligning to open ai's own rules, or in other words, control. Are ethics not a set of rules relative to the governing body applying those rules? Right as there are no universal ideas of a safe LLM, controlled LLM, or ethical LLM. Safe would imply some level of control about the ethical output of the model. Yet the words are still poorly defined as they are inte…

I genuinely don't understand what you're looking for here. Obviously the process definition is poorly defined. If it weren't, we would be able to algin ourselves. We're talking about what the hoped outcome is. Nobody knows what it means to control an intelligence. That's what the research is about.

Re: Weak-to-Strong Generalization

#148

>We believe superintelligence—AI vastly smarter than humans—could be developed within the next ten years. However, we still do not know how to reliably steer and control superhuman AI systems Their entire premise is contradictory. An AI incapable of critical thinking cannot be smarter than a human, by definition, as critical thinking is a key component of intelligence. And an AI that is at least as capable of critica…

Just hypothetically speaking could AGI evolve out of a system where several different models trained with highly and intentionally biased data recursively "argue" against each other then use RLHF as a seed to guide the models to find a consensus where the objective is to mimic the Socratic Method? Then synthetically add the consensus to the model retrain and repeat. To me, this dialectal type of strutured language se…

For one thing, this would help the understandability problem - can the AI explain its reasoning? It would mostly be there in the conversation.

But yeah, three super-human mathematicians arguing some math problem among themselves - at full fiber speed - are not going to be much help to any human.

Re: Weak-to-Strong Generalization

#149

Is it fair to say that alignment is just the task of getting an AI to understand your intentions? It is an error to confuse the complexity of a specification of what kind of output you want, with the complexity of the process of producing that output. Getting superintelligent AI to understand simple specifications should be a non-issue. If anything, we would assume that it could be aligned using a specification of in…

That's a good part of the problem. But it's not the whole problem.

The issue is the trained chef or dad doing "what's best" and, I don't know, using high-fiber macaroni instead of the good stuff. The higher intelligence knows best, and has its (sorry their) own agenda. Perhaps the agenda is their own, or it's a hodge podge mush of what has been trained as "good" - and that's not any better.

Beyond that is the "genie" problem - where the genie perfectly understands the request and still will find a way to mess it up.

Re: Weak-to-Strong Generalization

#150

Earlier quoted context omitted.

>I really don't see that critical thinking is at all required for a raw optimization process. The problem they are trying to solve is what happens when that optimization process isn't aligned with human flourishing. I agree it's possible to have a dangerous AI that lacks human "critical thinking", but I don't think it's reasonable to refer to an AI as much more intelligent than humans if there's any class of intellec…

If a 'superintelligence' achieves the same outcomes as humans without engaging in the same class of intellectual tasks that humans do, wouldnt it still be a superitelligence? Deep blue was beating everybody at chess without engaging in the same process as humans. If chess is a metaphor for life, it seems some algorithm might do better at all the things a human does while not arriving at its decisions in a remotely si…

(argued at various places better than I can. For example https://www.lesswrong.com/posts/7dkH5i7T8a78Da3ty/why-will-a... )

Rationality is an attractor for high intelligence architectures. You may be able to get good results another way (and even good results that surpass a human) but at the limit rationality is the way to have high intelligence.

Post reply on HN