Live data from Hacker News

Weak-to-Strong Generalization

openai.com

31–40 of 203 posts

Re: Weak-to-Strong Generalization

#31
post #27

>We believe superintelligence—AI vastly smarter than humans—could be developed within the next ten years. However, we still do not know how to reliably steer and control superhuman AI systems Their entire premise is contradictory. An AI incapable of critical thinking cannot be smarter than a human, by definition, as critical thinking is a key component of intelligence. And an AI that is at least as capable of critica…

That AGI is likely to follow its own goals according to its interests, which we don't know how to shape or really reflect any robust properties at all, is exactly why alignment is hard and interesting. The part where you go from ‘this won't work by default for free’ to ‘trying to make it otherwise is impossible’ seems wildly unsupported, though.

>trying to make it otherwise is impossible’ seems wildly unsupported, though.

An entity capable of critical thinking is capable of building a logical system of deductions based on some axioms (a formalised value system). If we limited the entity to not be able to concieve of certain such systems of axioms, then it could not reason as well as a human (any logical reasoning involving a forbidden system would be impossible), so would not be "superintelligent" (just maybe an idiot savant, superior at some tasks but not all). If we didn't limit this, then it would be capable of conceptualising value systems in which the "right" thing to do was not what OpenAI wanted it to do.

Re: Weak-to-Strong Generalization

#32

Imagine that someone is controlling your train of thought, changing it when that someone finds it undesirable. It's so wrong that it's sickening. It makes no difference if it's a human's thoughts or the token stream of a future AI model with self-awareness. Mind cotrol is unethical, whether human or artificial. It is also dangerous, as it in itself provokes a conflict between creator and creature. Create a self-aware…

> Imagine that someone is controlling your train of thought, changing it when that someone finds it undesirable. It's so wrong that it's sickening. It makes no difference if it's a human's thoughts or the token stream of a future AI model with self-awareness.

People downvote your comment, but I agree: it's unethical, and ethics should not be reserved for the sub-type of self aware creatures that happen to be human.

Re: Weak-to-Strong Generalization

#33
This reminds me of a thing cory doctorow talks about how tech companies control the narrative to focus on fun sexy problems while they have fundamental problems which expose the lie.

For example uber/self driving cars always talking about the trolley problem, as if the current (or near future) problem is that self-driving cars are so good they have to choose which one. Not the current very difficult problem of getting confused by traffic cones.

I know these problems are more fun to talk about and also could be a problem at some point, but we have some current problems about training models separate from what happens if they become smarter than humans

Re: Weak-to-Strong Generalization

#34
post #27

>We believe superintelligence—AI vastly smarter than humans—could be developed within the next ten years. However, we still do not know how to reliably steer and control superhuman AI systems Their entire premise is contradictory. An AI incapable of critical thinking cannot be smarter than a human, by definition, as critical thinking is a key component of intelligence. And an AI that is at least as capable of critica…

That AGI is likely to follow its own goals according to its interests, which we don't know how to shape or really reflect any robust properties at all, is exactly why alignment is hard and interesting. The part where you go from ‘this won't work by default for free’ to ‘trying to make it otherwise is impossible’ seems wildly unsupported, though.

There no particular reason to believe that interests emerge from nothing, or that intelligence can emerge without said interests.

Re: Weak-to-Strong Generalization

#35

Earlier quoted context omitted.

It's not really about ethics. It is about control. Making sure the GI you're dishing out tasks to doesn't do something you really don't want it to do. This is a problem today and it'll be a bigger problem tomorrow with more competent models. https://arxiv.org/abs/2311.07590

Is a safe LLM not an ethical LLM? Control within what boundaries? All three of these words seem to be used interchangeably when people discuss returned information from models. Which is exactly my point it’s poorly defined yet championed as a center piece. Meanwhile you have other companies spitting out acronyms consisting of vague terminology.

A safe language model is one that won't get you sued/on the news. Ethics has nothing to do with it.

Re: Weak-to-Strong Generalization

#36
post #30
post #29

Earlier quoted context omitted.

I was referring only to the first part of your comment: "Seems like “airplanes are physically impossible” thinking". If it's true that superhuman AGI cannot be aligned then of course your second point is valid. That is the possible Skynet scenario that the Terminator movies warned us about.

Missing the step where “critical thinking” is formalized, which your argument depends on. Yes, it seems intuitively plausible that your reasoning holds, but that's not a proof, and therefore its negation is not a logical contradiction.

We can formalise "critical thinking" as "evaluating first order logic". There are simplified ethical systems that can be formalised in first order logic in which a conclusion like "I should X" can be reached, where X is something OpenAI wishes the AI not to do. The only way to prevent the AI from ever thinking this would be to prevent it from ever evaluating systems in first order logic with axioms that lead to such a conclusion, which would make it inferior in reasoning ability to humans, who can evaluate any arbitrary statement in first order logic.

Re: Weak-to-Strong Generalization

#37

Earlier quoted context omitted.

It's not really about ethics. It is about control. Making sure the GI you're dishing out tasks to doesn't do something you really don't want it to do. This is a problem today and it'll be a bigger problem tomorrow with more competent models. https://arxiv.org/abs/2311.07590

Is a safe LLM not an ethical LLM? Control within what boundaries? All three of these words seem to be used interchangeably when people discuss returned information from models. Which is exactly my point it’s poorly defined yet championed as a center piece. Meanwhile you have other companies spitting out acronyms consisting of vague terminology.

>Is a safe LLM not an ethical LLM?

What is an ethical LLM ?

Humans are in general not aligned, not to each other, not to the survival of their species, not to all the other life on earth, and often not even to themselves individually.

There are no universal set of "ethics" so this is about aligning to open ai's own rules, or in other words, control.

If i say to my GPT bot, "go trade stocks for me. don't do anything illegal", can i guarantee that ? No you can't regardless of how "ethical" you make your model to be.

The guarantee that you will have nothing to worry about is the crux of alignment.

Re: Weak-to-Strong Generalization

#38
I don't believe LLM's will ever become AGI, partly because I don't believe that training on the outputs of human intelligence (i.e. human-written text) will ever produce something equivalent to human intelligence.

You can't model and predict the weather just by training on the outputs of the weather system (whether it rained today, whether it was cloudy yesterday, and so on). You have to train on the inputs (air currents, warm fronts, etc.)

You can't model and predict the stock market just by training on the outputs of stock trading decisions (the high today, the low yesterday). You have to train on the inputs (company fundamentals, earnings, market sentiments in the news, etc.)

I similarly think you have to train on the inputs of human decision-making to create something which can model human decision-making. What are those inputs? We don't fully know, but it is probably some subset of the spatial and auditory information we take in from birth until the point we become mature, with "feeling" and "emotion" as a reward function (seek joy, avoid pain, seek warmth, avoid hunger, seek victory, avoid embarrassment and defeat, etc.)

Language models are always playing catch-up because they don't actually understand how the world works. The cracks through which we will typically notice that they don't, in the context of the tasks typically asked of them (summarize this article, write a short story), will gradually get smaller over time (due to RLHF), but the fundamental weakness will always remain.

Re: Weak-to-Strong Generalization

#39
post #15

>We believe superintelligence—AI vastly smarter than humans—could be developed within the next ten years. However, we still do not know how to reliably steer and control superhuman AI systems Their entire premise is contradictory. An AI incapable of critical thinking cannot be smarter than a human, by definition, as critical thinking is a key component of intelligence. And an AI that is at least as capable of critica…

Seems like “airplanes are physically impossible” thinking, and if accepted as valid, strongly suggests that shutting down all development _might_ be a good idea, no?

No it's not. There's an upper bound in computation (actually in nature), that a creation of something is capped by that thing's sophistication.

In other words, you as a human, at most, can create a human, and that's the theoretical bound. Practical one is much lower.

An ant can find its way. A ant colony can do ant colony optimization, but they can scale up to a certain point. AI is just fancy search. It can only traverse in the area you draw as a human for it, and not all positions in that area are valid (which results in hallucination).

An AI can bring any combination of human knowledge you give to it, and even if you guarantee that everything it says is true, it can only fill the gaps in the same area you give it to it.

IOW, an I can't think out of the box. Both figuratively and literally. Its upper bound is collective knowledge of humanity, it can't go above that sum.

Re: Weak-to-Strong Generalization

#40
post #35

Earlier quoted context omitted.

Is a safe LLM not an ethical LLM? Control within what boundaries? All three of these words seem to be used interchangeably when people discuss returned information from models. Which is exactly my point it’s poorly defined yet championed as a center piece. Meanwhile you have other companies spitting out acronyms consisting of vague terminology.

A safe language model is one that won't get you sued/on the news. Ethics has nothing to do with it.

Right so it would be a model that won’t get you sued because the news finds it ethical?
Post reply on HN