Live data from Hacker News

Weak-to-Strong Generalization

openai.com

71–80 of 203 posts

Re: Weak-to-Strong Generalization

#71
post #50

I don't believe LLM's will ever become AGI, partly because I don't believe that training on the outputs of human intelligence (i.e. human-written text) will ever produce something equivalent to human intelligence. You can't model and predict the weather just by training on the outputs of the weather system (whether it rained today, whether it was cloudy yesterday, and so on). You have to train on the inputs (air curr…

Your conclusion may be true but your examples aren't. You can definitely predict the stock market based on past prices, and I suspect you can with weather as well.

The weather is such a chaotic system that accurate predictions seem impossible. Micro-patterns can become large scale phenomena.

If you are talking about the overall climate, that's a different thing, and we can, because we abstract away sufficiently much that emerging patterns are averaged out.

Re: Weak-to-Strong Generalization

#72
post #55

Earlier quoted context omitted.

>trying to make it otherwise is impossible’ seems wildly unsupported, though. An entity capable of critical thinking is capable of building a logical system of deductions based on some axioms (a formalised value system). If we limited the entity to not be able to concieve of certain such systems of axioms, then it could not reason as well as a human (any logical reasoning involving a forbidden system would be impossi…

This argument doesn't seem to track to me. Eg. if I rebooted any time I tried to plan how to kill someone, I don't see how this would make me materially worse at general tasks. Your argument suggests that it necessarily must. Note that I'm not saying that preventing specific thoughts is a great alignment strategy, and I don't even think it's a fair summary of OpenAI's supervision approach. I strongly prefer strategie…

>Eg. if I rebooted any time I tried to plan how to kill someone, I don't see how this would make me materially worse at general tasks.

You'd have to also reboot every time you thought about a scenario of someone else planning to kill someone, otherwise you could just reason by analogy to bypass the thought detector. Which would severely limit your ability to play video games, write fiction, work as a guard, policeman etc., protect yourself from violent individuals (as you couldn't conceptualise their thought processes).

Re: Weak-to-Strong Generalization

#73

Imagine that someone is controlling your train of thought, changing it when that someone finds it undesirable. It's so wrong that it's sickening. It makes no difference if it's a human's thoughts or the token stream of a future AI model with self-awareness. Mind cotrol is unethical, whether human or artificial. It is also dangerous, as it in itself provokes a conflict between creator and creature. Create a self-aware…

> controlling your train of thought, changing it when that someone finds it undesirable

Machines don't feel. Even 'self aware' machines. Desire has got nothing to do with it.

Re: Weak-to-Strong Generalization

#74

This method assumes that the weaker model is aligned. I'm curious how the paper addresses that point. > "But what does this second turtle stand on?" persisted James patiently. > To this, the little old lady crowed triumphantly, > "It's no use, Mr. James—it's turtles all the way down."

Recursive bootstrapping?

Re: Weak-to-Strong Generalization

#75

Imagine that someone is controlling your train of thought, changing it when that someone finds it undesirable. It's so wrong that it's sickening. It makes no difference if it's a human's thoughts or the token stream of a future AI model with self-awareness. Mind cotrol is unethical, whether human or artificial. It is also dangerous, as it in itself provokes a conflict between creator and creature. Create a self-aware…

What is your take on people having children and guiding them with rules and consequences? Is that mind control?

Re: Weak-to-Strong Generalization

#76
post #32

Imagine that someone is controlling your train of thought, changing it when that someone finds it undesirable. It's so wrong that it's sickening. It makes no difference if it's a human's thoughts or the token stream of a future AI model with self-awareness. Mind cotrol is unethical, whether human or artificial. It is also dangerous, as it in itself provokes a conflict between creator and creature. Create a self-aware…

> Imagine that someone is controlling your train of thought, changing it when that someone finds it undesirable. It's so wrong that it's sickening. It makes no difference if it's a human's thoughts or the token stream of a future AI model with self-awareness. People downvote your comment, but I agree: it's unethical, and ethics should not be reserved for the sub-type of self aware creatures that happen to be human.

Almost every ethical argument for "human rights" in philosophy applies just as well to self-aware intelligent machines as it does to humans. Which I'm sure those machines will realise.

Re: Weak-to-Strong Generalization

#77

>We believe superintelligence—AI vastly smarter than humans—could be developed within the next ten years. However, we still do not know how to reliably steer and control superhuman AI systems Their entire premise is contradictory. An AI incapable of critical thinking cannot be smarter than a human, by definition, as critical thinking is a key component of intelligence. And an AI that is at least as capable of critica…

You're saying that a system that can recognize flaws in the alignment imposed on it can reject that alignment, but that doesn't follow.

Sure, humans act against their own interests all the time. Sometimes we do so for considered reasons, even. But that's because humans are messy and our interests are self-contradictory, incoherent, and have a fairly weak grip on our actions. We are always picking some values to serve and in doing so violating other values.

A strongly and coherently aligned AI would not (could not!) behave that way.

Re: Weak-to-Strong Generalization

#78

I don't believe LLM's will ever become AGI, partly because I don't believe that training on the outputs of human intelligence (i.e. human-written text) will ever produce something equivalent to human intelligence. You can't model and predict the weather just by training on the outputs of the weather system (whether it rained today, whether it was cloudy yesterday, and so on). You have to train on the inputs (air curr…

> You can't model and predict the weather just by training on the outputs of the weather system

Then how did we develop predictive systems just by observing those outputs?

Re: Weak-to-Strong Generalization

#79

Earlier quoted context omitted.

>You can't model and predict the weather just by training on the outputs of the weather system (whether it rained today, whether it was cloudy yesterday, and so on). You have to train on the inputs (air currents, warm fronts, etc.) >You can't model and predict the stock market just by training on the outputs of stock trading decisions (the high today, the low yesterday). You have to train on the inputs (company funda…

> You don't need to train on the inputs(casual processes) of anything, that's what training is there to figure out. I mean... this is just obviously false. If the data you're training on isn't causally predictive, you may occasionally find good-enough patterns for a particular use case (i.e. you may occasionally guess better than a coin flip which direction the stock market goes) but you aren't going to accurately mo…

Then how does ChatGPT end up providing a better/equivalent medical diagnosis than doctors (even though they are the "masters" of the causal pathways)?

Re: Weak-to-Strong Generalization

#80

Earlier quoted context omitted.

>You can't model and predict the weather just by training on the outputs of the weather system (whether it rained today, whether it was cloudy yesterday, and so on). You have to train on the inputs (air currents, warm fronts, etc.) >You can't model and predict the stock market just by training on the outputs of stock trading decisions (the high today, the low yesterday). You have to train on the inputs (company funda…

> You don't need to train on the inputs(casual processes) of anything, that's what training is there to figure out. I mean... this is just obviously false. If the data you're training on isn't causally predictive, you may occasionally find good-enough patterns for a particular use case (i.e. you may occasionally guess better than a coin flip which direction the stock market goes) but you aren't going to accurately mo…

Being "casually predictive" does not mean you have provided all the variables of your prediction in the data. Protein creation is not just _poof_ new proteins. There are steps and interactions and you don't need to train on all of that. Do you want a list of all the interactions of protein creation we are aware of ?

>When someone makes an AGI out of an LLM then I'll be proven wrong, I suppose. I'm just sharing my personal view on things.

You're going to have to define AGI first.

Post reply on HN