Live data from Hacker News

Weak-to-Strong Generalization

openai.com

11–20 of 203 posts

Re: Weak-to-Strong Generalization

#11

> Figuring out how to align future superhuman AI systems to be safe has never been more important They love using the word “safe” and I’m pretty sure it’s 99% PR, because reading their other “papers” on Safety & Alignment seems to not really identify or define safety bounds at all. You’d think this has something to do with ethics but we all know there are no longer any ethically concerned leaders at their workplace.…

We're deliberately trying to create something with the capability to also create. It's not ridiculous to be concerned about what we might end up with.

Re: Weak-to-Strong Generalization

#12

> Figuring out how to align future superhuman AI systems to be safe has never been more important They love using the word “safe” and I’m pretty sure it’s 99% PR, because reading their other “papers” on Safety & Alignment seems to not really identify or define safety bounds at all. You’d think this has something to do with ethics but we all know there are no longer any ethically concerned leaders at their workplace.…

It's not really about ethics. It is about control. Making sure the GI you're dishing out tasks to doesn't do something you really don't want it to do.

This is a problem today and it'll be a bigger problem tomorrow with more competent models. https://arxiv.org/abs/2311.07590

Re: Weak-to-Strong Generalization

#13
post #5
post #2

I hope OpenAI will continue to prioritize working on these crucial questions after the boardroom drama.

Weren't all the board members who wanted to prioritize these crucial questions fired? Hopefully the employees who were hired to perform this work have enough inertia to continue until the board recovers (if it recovers).

They got to choose their replacements; it's not like they were forced out to be replaced by anybody Altman wanted.

Re: Weak-to-Strong Generalization

#14
This method assumes that the weaker model is aligned. I'm curious how the paper addresses that point.

> "But what does this second turtle stand on?" persisted James patiently.

> To this, the little old lady crowed triumphantly,

> "It's no use, Mr. James—it's turtles all the way down."

Re: Weak-to-Strong Generalization

#15

>We believe superintelligence—AI vastly smarter than humans—could be developed within the next ten years. However, we still do not know how to reliably steer and control superhuman AI systems Their entire premise is contradictory. An AI incapable of critical thinking cannot be smarter than a human, by definition, as critical thinking is a key component of intelligence. And an AI that is at least as capable of critica…

Seems like “airplanes are physically impossible” thinking, and if accepted as valid, strongly suggests that shutting down all development _might_ be a good idea, no?

Re: Weak-to-Strong Generalization

#16
post #13
post #5

Earlier quoted context omitted.

Weren't all the board members who wanted to prioritize these crucial questions fired? Hopefully the employees who were hired to perform this work have enough inertia to continue until the board recovers (if it recovers).

They got to choose their replacements; it's not like they were forced out to be replaced by anybody Altman wanted.

Oh! Thank you for the update. I had missed that.

Re: Weak-to-Strong Generalization

#17

>We believe superintelligence—AI vastly smarter than humans—could be developed within the next ten years. However, we still do not know how to reliably steer and control superhuman AI systems Their entire premise is contradictory. An AI incapable of critical thinking cannot be smarter than a human, by definition, as critical thinking is a key component of intelligence. And an AI that is at least as capable of critica…

I think the premise is dubious as well but since they are deadset on creating this intelligence, they might as well try to figure out a way to control it, hopeless as it may seem.

Re: Weak-to-Strong Generalization

#18

> Figuring out how to align future superhuman AI systems to be safe has never been more important They love using the word “safe” and I’m pretty sure it’s 99% PR, because reading their other “papers” on Safety & Alignment seems to not really identify or define safety bounds at all. You’d think this has something to do with ethics but we all know there are no longer any ethically concerned leaders at their workplace.…

It's not really about ethics. It is about control. Making sure the GI you're dishing out tasks to doesn't do something you really don't want it to do. This is a problem today and it'll be a bigger problem tomorrow with more competent models. https://arxiv.org/abs/2311.07590

Is a safe LLM not an ethical LLM? Control within what boundaries? All three of these words seem to be used interchangeably when people discuss returned information from models. Which is exactly my point it’s poorly defined yet championed as a center piece. Meanwhile you have other companies spitting out acronyms consisting of vague terminology.

Re: Weak-to-Strong Generalization

#19

>We believe superintelligence—AI vastly smarter than humans—could be developed within the next ten years. However, we still do not know how to reliably steer and control superhuman AI systems Their entire premise is contradictory. An AI incapable of critical thinking cannot be smarter than a human, by definition, as critical thinking is a key component of intelligence. And an AI that is at least as capable of critica…

Yes the superhuman AI would need to be coerced to remain politically correct. How can we coerce an AI?

Re: Weak-to-Strong Generalization

#20
post #15

>We believe superintelligence—AI vastly smarter than humans—could be developed within the next ten years. However, we still do not know how to reliably steer and control superhuman AI systems Their entire premise is contradictory. An AI incapable of critical thinking cannot be smarter than a human, by definition, as critical thinking is a key component of intelligence. And an AI that is at least as capable of critica…

Seems like “airplanes are physically impossible” thinking, and if accepted as valid, strongly suggests that shutting down all development _might_ be a good idea, no?

No. This is a logical contradiction.

Edit: I mean the comment you are replying to is showing there is a logical contradiction.

If the AI is capable of critical thinking then it will independently form its own judgements and conclusions. If it simply believes whatever we tell it to believe, then that is not critical thinking, by definition.

Post reply on HN