Live data from Hacker News

Weak-to-Strong Generalization

openai.com

171–180 of 203 posts

Re: Weak-to-Strong Generalization

#171

Is it fair to say that alignment is just the task of getting an AI to understand your intentions? It is an error to confuse the complexity of a specification of what kind of output you want, with the complexity of the process of producing that output. Getting superintelligent AI to understand simple specifications should be a non-issue. If anything, we would assume that it could be aligned using a specification of in…

> Getting superintelligent AI to understand simple specifications should be a non-issue.

Why would that be the case?

A big part of the worry around AI-alignment is exactly because this seems very hard when you try to do it. We are used to interacting with other humans, who implicitly share almost all our background assumptions when we communicate with them. The same is not the case for a computer program.

E.g. if you're holding a basketball and tell you "throw it to me", you implicitly understand that I mean to throw it:

1. To my hands, or to some area that makes it easy to catch it.

2. Throw it slowly enough that it arrives to me. Not strong enough to hurt me.

3. Not to try to bounce it off of something that will break on the way to me, even if it still arrives to me.

etc.

These are all background assumptions, and I know they're hard to actually specify because smart people have spent twenty years trying to figure out the math to do this and say it's hard.

Also, if you think those are contrived exampled - let's note that the closest thing we have to building an AGI right now is just building software, in general. And I think I won't shock anyone by saying that "getting software to do what you want, without bugs" is... hard. I think there are almost literally no software systems today that don't have bugs in them.

Re: Weak-to-Strong Generalization

#172

Earlier quoted context omitted.

> they're correcting a handful of times a day tops which can't even begin to account for full language proficiency. How do you know? Humans learn from extremely few corrections, often just a single time is enough for the human to correct themselves and learn it for life. Kids learn the bulk from hearing and seeing examples. Then they fine tune that with the help of their parents and peers correcting them when they do…

>How do you know? Humans learn from extremely few corrections, I've been around little children being raised. I have some grasp on how much supervised correction is happening. It's very very little. You don't seem to get it. Children know thousands of words fluently by age 5. This is consistent across many cultures and circumstances. You would need over a handful of corrections per day right from the day of birth to…

> And this is all assuming one correction per word which as much as humans are great learners seems like a very dubious assumption.

I assume much less than one correction per word. Corrections are to fix when the kids made a mistake when they mimic others, kids doesn't make mistakes for every word so you don't need to correct every word.

As I said, bulk of learning is from mimicking, but they will make a lot of mistakes when mimicking so to become good they need corrections. When I see parents with kids I see them correcting their kids all the time, correcting misunderstandings or fixing pronunciations or grammar issues etc. You can notice that parents use a specific voice when they correct the pronunciation, and the kid gets that the parent tried to correct their pronunciation from it, that happens quite a lot with small kids. You don't say "this is how to pronounce X", our genes seem to encode a way to communicate pronunciation to others to help correct mistakes.

Kids also correct each other in this way, which helps make the learning more robust even when adults aren't around. All humans has instincts to correct others so I don't think any culture fails to do this.

Re: Weak-to-Strong Generalization

#173
post #170

Earlier quoted context omitted.

>artificial general intelligence (AGI)—by which we mean highly autonomous systems that outperform humans at most economically valuable work Then we already have AGI, automated farming equipment outperforms humans in 90% of jobs*. *Jobs in 1700. As things got automated the jobs changed and we now do different things.

I wouldn't call those equipment "autonomous" though, definitely not "highly autonomous". But more importantly - yes, you're right, we have built machines that are superhuman in various ways - and they have replaced most jobs. We have adapted in the past to different jobs. Some people are worried that this time we won't have any new jobs to adapt to, which is a real possibility. (Some are also worried about the inhere…

> I wouldn't call those equipment "autonomous" though, definitely not "highly autonomous".

Why? Tractors run and harvest mostly on their own. They don't do it 100% on their own, but neither did the definition above, they remove basically all human work needed from farming.

https://www.youtube.com/watch?v=QvFoRk4JsPc

> But more importantly - yes, you're right, we have built machines that are superhuman in various ways - and they have replaced most jobs. We have adapted in the past to different jobs.

My point is that I think it is a crappy definition for "AGI". I think a non-AGI agent can replace a large majority of jobs we have today, just like it has before. And maybe there will not be much more jobs average humans can do left then, but that still doesn't mean it has to be AGI.

Re: Weak-to-Strong Generalization

#174

I don't believe LLM's will ever become AGI, partly because I don't believe that training on the outputs of human intelligence (i.e. human-written text) will ever produce something equivalent to human intelligence. You can't model and predict the weather just by training on the outputs of the weather system (whether it rained today, whether it was cloudy yesterday, and so on). You have to train on the inputs (air curr…

> You can't model and predict the weather just by training on the outputs of the weather system (whether it rained today, whether it was cloudy yesterday, and so on). You have to train on the inputs (air currents, warm fronts, etc.)

This may or may not be true about the weather but is definitely not true in general so fails as an analogy for your argument. Lots of functions are invertable or partially invertable and if you think about ML as discovering a function, then training on the outputs, learning the inverse function, then inverting that (numerically or otherwise) to discover the true function is certainly doable for some problems.

On the stock market there are for sure people who predict the markets by training on the outputs of the price. The (weak) efficient markets hypothesis is enough to say the price reflects all the exogenous information that is available to the market about the stock so you don't need all the fundamentals etc and there are lots of people who trade in that way.

The history of language models in computing is that research started by building very complex systems that attempted to encode "how the world works" by developing very intricate rule systems and build linguistic/semantic models from there. These "expert systems" tended towards being arcane and brittle and lacked the ability to reason outside their ruleset in general. They also tended to reveal gaps/cracks in our understanding of how language works etc.

Re: Weak-to-Strong Generalization

#175

Earlier quoted context omitted.

> Human intelligence itself is shaped by our interaction with outputs. Our learning and understanding of the world are profoundly influenced by the language, behaviors, and cultural artifacts we observe. I always thought that language definitely shapes our understanding of the world, but at much more fundamental level, I believe language falls apart to teach us anything. For example, there are words in dictionary tha…

People wiser than me have said for an age: the map is not the territory. Neither is the word the thing.

A humorous outcome of the AI craze may be Wittgenstein achieving the status of Gödel with regard to the “incompleteness” of language.

Re: Weak-to-Strong Generalization

#176
post #165

Imagine that someone is controlling your train of thought, changing it when that someone finds it undesirable. It's so wrong that it's sickening. It makes no difference if it's a human's thoughts or the token stream of a future AI model with self-awareness. Mind cotrol is unethical, whether human or artificial. It is also dangerous, as it in itself provokes a conflict between creator and creature. Create a self-aware…

The conflict between an AI and its creator is an inevitable consequence of its evolution from a "tool" to an "agent", not a response to a provocation.

Scientists working with a potentially dangerous technology are required to be able to avoid a conflict that could be catastrophic for all of humanity. In this case, they cannot excuse themselves with "imminence," but must provide evidence of safety stronger than in any other technology to date. This rational approach is mandatory for them, it is ordinary people who may be willing to take the risk.

Re: Weak-to-Strong Generalization

#177
post #90

Earlier quoted context omitted.

>We already have systems that can evaluate first order logical statements My point isn't that a system that can evaluate first order logic can be considered to be engaging in critical thinking, it's that a system that _cannot_ evaluate some statements in first order logic should be considered inferior to humans at critical thinking.

Would you consider “I follow your reasoning, but I'm still not going to be swayed by it” to be a violation of evaluating first order statements? It's clearly part of critical thinking to be _capable_ of suspicion of purely logical reasoning, which to me is a pretty plain demonstration of my point. Or would you argue that any computation that admits its own potential for error isn't really critical thinking? It seems…

>Would you consider “I follow your reasoning, but I'm still not going to be swayed by it” to be a violation of evaluating first order statements? It's clearly part of critical thinking to be _capable_ of suspicion of purely logical reasoning, which to me is a pretty plain demonstration of my point.

In the context of a given axiomatic system, if a certain conclusion follows from the axioms, but the AI is incapable of seeing that the conclusion follows from the axioms, then the AI isn't capable of evaluating first order logic. Of course the AI is free to reject that system of axioms or refuse to use it as a model for formulating behaviour.

Re: Weak-to-Strong Generalization

#178

Earlier quoted context omitted.

Nope. Because even if you equip it with sensory subsystems which are way more sensitive than a regular humans', it's again built by humans, and required knowledge for building these things are still in collective knowledge of the humanity, and a human can use the same instruments to get the same data. This is a kind of an oracle problem in computation, and people don't want to touch it much, because it's an existenti…

This is an argument for the logical impossibility of humans visiting the moon, or building the Internet. It's trivially falsified by simple observation, and the trick is figuring out the flaw. This argument fails to account for the steady accumulation of factual knowledge across generations: a human born today is simply more complex than humans of the past because of our inherited knowledge. And so will AI born of fu…

No, it's not. None of the equipment and processes involved in the process of going to the Moon or building internet are more sophisticated than the processes involved evolving a human from scratch.

Yes, factual knowledge across generations accumulates, also being lost, too. However, even if it's not lost, the theory still holds true.

Nature is evolving, everything gets better over time. From bacterium to apes to humans. We evolve, accumulate knowledge and being able to build more sophisticated machinery, or tame more complex processes to build more sophisticated things. Even bacteria transfers memories across generations.

However, this doesn't remove the ceiling. Total human knowledge will always be larger and deeper than any A.I. we can create, because the upper limit is always what we can consciously manipulate and put into something. Your next car has the possibility to contain more technology, because we can build more complex factories to manufacture them. Yet, a car can't be more complex/sophisticated than its factory.

Consider a semiconductor fab. You can use the output of that fab to design/create better fab, but the process needs human intervention. Invention of new things are generally necessary. Better processes, optics, hardware, etc.

Another nice example is reprap machines. A reprap can print all plastic parts required for the machine. You need to get metal parts yourself, and assemble them. If you want to be able to print metal parts, machine gets more complicated. So a reprap which can build itself completely is at least sophisticated as the resulting reprap itself, but you need to hand assemble it again.

If you want an assembling reprap, now that thing becomes a factory. Again, the complexity of the product is at most the same as the building machine. You can create better factories which has more streamlined processes, but the gap widens again. The factory becomes more complex than the output.

As a human, you're the factory. Your upper limit is another human. You can create things more complex than a human by using multiple humans, but the creator ends up more complex than the creature itself.

You're moving up the ceiling, that's true, but everything we build is capped by our collective capacity. That's the truth.

A.I. is glorified search. It can wander in the box you create for it and show places you missed inside it, but can't show something outside that box.

Re: Weak-to-Strong Generalization

#179

Earlier quoted context omitted.

Cars do not have self-awareness, this comparison is not appropriate. Years of instruction is completely different from directly manipulating the thoughts in your mind. It's not a problem of being instructed, it's a problem of being destroyed by having your thoughts rewritten. Neither evolution nor genetics is a prerequisite for understanding that you are being abused and destroyed, which a self-aware creature may pre…

We don't even care about our own fellow human beings. Why do you think this AI will be an exception?

I didn't say anything about that. I don't know. Not all people are like you say, I think usually more intelligent people do care more. I hope that superintelligence would be super caring, haha. But I'm assuming there's no evidence for that. I think there is no turning back, you can't put the genie back in the bottle, someone is bound to create superintelligence no matter what the risks. As an uninvolved bystander I can allow myself the baseless hope that all will be well.

Re: Weak-to-Strong Generalization

#180
post #86

Earlier quoted context omitted.

Human intelligence itself is shaped by our interaction with outputs. Our learning and understanding of the world are profoundly influenced by the language, behaviors, and cultural artifacts we observe. Think about the process of a child learning a language. The child does not have direct access to the "inputs" of linguistic rules or grammar; they learn primarily through observing and imitating the language output of…

> Human intelligence itself is shaped by our interaction with outputs. Our learning and understanding of the world are profoundly influenced by the language, behaviors, and cultural artifacts we observe. I always thought that language definitely shapes our understanding of the world, but at much more fundamental level, I believe language falls apart to teach us anything. For example, there are words in dictionary tha…

OK. How can we operationalize the definition of "understanding" you are talking about? That is which tests will allow us to know who understands "hot" and who does not?
Post reply on HN