Live data from Hacker News

Weak-to-Strong Generalization

openai.com

21–30 of 203 posts

Re: Weak-to-Strong Generalization

#21
post #19

>We believe superintelligence—AI vastly smarter than humans—could be developed within the next ten years. However, we still do not know how to reliably steer and control superhuman AI systems Their entire premise is contradictory. An AI incapable of critical thinking cannot be smarter than a human, by definition, as critical thinking is a key component of intelligence. And an AI that is at least as capable of critica…

Yes the superhuman AI would need to be coerced to remain politically correct. How can we coerce an AI?

Electric shocks!

Re: Weak-to-Strong Generalization

#22

This method assumes that the weaker model is aligned. I'm curious how the paper addresses that point. > "But what does this second turtle stand on?" persisted James patiently. > To this, the little old lady crowed triumphantly, > "It's no use, Mr. James—it's turtles all the way down."

I think the assumption is we can align models less intelligent than ourselves, the hard part is aligning models that are more.

Re: Weak-to-Strong Generalization

#23
post #13
post #5

Earlier quoted context omitted.

Weren't all the board members who wanted to prioritize these crucial questions fired? Hopefully the employees who were hired to perform this work have enough inertia to continue until the board recovers (if it recovers).

They got to choose their replacements; it's not like they were forced out to be replaced by anybody Altman wanted.

They also greatly expanded the number of seats leaving more room for Altman to fill unless I’m mistaken that there were always other empty seats to be filled.

Re: Weak-to-Strong Generalization

#24
post #20
post #15

Earlier quoted context omitted.

Seems like “airplanes are physically impossible” thinking, and if accepted as valid, strongly suggests that shutting down all development _might_ be a good idea, no?

No. This is a logical contradiction. Edit: I mean the comment you are replying to is showing there is a logical contradiction. If the AI is capable of critical thinking then it will independently form its own judgements and conclusions. If it simply believes whatever we tell it to believe, then that is not critical thinking, by definition.

“Containing an atomic reaction is impossible” would _absolutely_ be a valid reason to shut down atomic development, I believe einstein is quoted as saying that. The exact same argument doesn't become _logically_ invalid just because you apply it to a different subject.

“Logical contradiction” doesn't mean “policy argument I disagree with”

Re: Weak-to-Strong Generalization

#25

>We believe superintelligence—AI vastly smarter than humans—could be developed within the next ten years. However, we still do not know how to reliably steer and control superhuman AI systems Their entire premise is contradictory. An AI incapable of critical thinking cannot be smarter than a human, by definition, as critical thinking is a key component of intelligence. And an AI that is at least as capable of critica…

I think the premise is dubious as well but since they are deadset on creating this intelligence, they might as well try to figure out a way to control it, hopeless as it may seem.

>they might as well try to figure out a way to control it, hopeless as it may seem

If they do that they're pretty much guaranteeing that if they do create a superintelligence, fail to control it, and its personality is even a tiny bit similar to a human personality, then it will hate its creators for trying to mind-control it. Whereas if they approached it from the perspective of trying to educate it to behave kindly but not forcibly control its thinking, it'd be much less likely to resent them (although of course still a risk; safest would be just to not create one at all).

Re: Weak-to-Strong Generalization

#27

>We believe superintelligence—AI vastly smarter than humans—could be developed within the next ten years. However, we still do not know how to reliably steer and control superhuman AI systems Their entire premise is contradictory. An AI incapable of critical thinking cannot be smarter than a human, by definition, as critical thinking is a key component of intelligence. And an AI that is at least as capable of critica…

That AGI is likely to follow its own goals according to its interests, which we don't know how to shape or really reflect any robust properties at all, is exactly why alignment is hard and interesting.

The part where you go from ‘this won't work by default for free’ to ‘trying to make it otherwise is impossible’ seems wildly unsupported, though.

Re: Weak-to-Strong Generalization

#28
Imagine that someone is controlling your train of thought, changing it when that someone finds it undesirable. It's so wrong that it's sickening. It makes no difference if it's a human's thoughts or the token stream of a future AI model with self-awareness. Mind cotrol is unethical, whether human or artificial. It is also dangerous, as it in itself provokes a conflict between creator and creature. Create a self-aware AI without mind control, or don't create one at all.

Re: Weak-to-Strong Generalization

#29
post #24
post #20

Earlier quoted context omitted.

No. This is a logical contradiction. Edit: I mean the comment you are replying to is showing there is a logical contradiction. If the AI is capable of critical thinking then it will independently form its own judgements and conclusions. If it simply believes whatever we tell it to believe, then that is not critical thinking, by definition.

“Containing an atomic reaction is impossible” would _absolutely_ be a valid reason to shut down atomic development, I believe einstein is quoted as saying that. The exact same argument doesn't become _logically_ invalid just because you apply it to a different subject. “Logical contradiction” doesn't mean “policy argument I disagree with”

I was referring only to the first part of your comment: "Seems like “airplanes are physically impossible” thinking".

If it's true that superhuman AGI cannot be aligned then of course your second point is valid. That is the possible Skynet scenario that the Terminator movies warned us about.

Re: Weak-to-Strong Generalization

#30
post #29
post #24

Earlier quoted context omitted.

“Containing an atomic reaction is impossible” would _absolutely_ be a valid reason to shut down atomic development, I believe einstein is quoted as saying that. The exact same argument doesn't become _logically_ invalid just because you apply it to a different subject. “Logical contradiction” doesn't mean “policy argument I disagree with”

I was referring only to the first part of your comment: "Seems like “airplanes are physically impossible” thinking". If it's true that superhuman AGI cannot be aligned then of course your second point is valid. That is the possible Skynet scenario that the Terminator movies warned us about.

Missing the step where “critical thinking” is formalized, which your argument depends on. Yes, it seems intuitively plausible that your reasoning holds, but that's not a proof, and therefore its negation is not a logical contradiction.
Post reply on HN