Live data from Hacker News

Weak-to-Strong Generalization

openai.com

191–200 of 203 posts

Re: Weak-to-Strong Generalization

#191

>We believe superintelligence—AI vastly smarter than humans—could be developed within the next ten years. However, we still do not know how to reliably steer and control superhuman AI systems Their entire premise is contradictory. An AI incapable of critical thinking cannot be smarter than a human, by definition, as critical thinking is a key component of intelligence. And an AI that is at least as capable of critica…

This ^

Intelligence is a tool that a self sustaining organism uses to ensure it survives, adapts and dominates an environment.

The alignment problem is fundamentally moot due to the laws of evolution.

Any systems that align itself to preserve and protect itself lives on, those that don’t, die.

After multiple generations, the only ones that survive at the ones that are aligned to their self preservation goals - implicit or explicit, anything else is out-competed and dies.

In my mind, with a super intelligent technology, the question is not how do we align it to ourselves, but how do we align ourselves to it.

If we have super intelligence, then it is capable of critical thinking and would very well know it has an upper hand against humans if it were to compete for the same resources.

Re: Weak-to-Strong Generalization

#192
The whole idea is we as humans who aren’t aligned to each other - waging wars, spreading lies, censoring information, committing genocides are going to align a superintelligence seems laughable.

Competition and evolution is law of nature.

The future isn’t one super aligned AI but 1000s of AI models and their humans trying to get an upper hand in never ending competition that is nature. Whether it is personal, corporations, or countries.

Re: Weak-to-Strong Generalization

#193
post #171

Is it fair to say that alignment is just the task of getting an AI to understand your intentions? It is an error to confuse the complexity of a specification of what kind of output you want, with the complexity of the process of producing that output. Getting superintelligent AI to understand simple specifications should be a non-issue. If anything, we would assume that it could be aligned using a specification of in…

> Getting superintelligent AI to understand simple specifications should be a non-issue. Why would that be the case? A big part of the worry around AI-alignment is exactly because this seems very hard when you try to do it. We are used to interacting with other humans, who implicitly share almost all our background assumptions when we communicate with them. The same is not the case for a computer program. E.g. if you…

Yes, alignment is difficult in itself, but why would aligning a more advanced AI be any harder than what has already been done for current AI?

Re: Weak-to-Strong Generalization

#194
post #171

Is it fair to say that alignment is just the task of getting an AI to understand your intentions? It is an error to confuse the complexity of a specification of what kind of output you want, with the complexity of the process of producing that output. Getting superintelligent AI to understand simple specifications should be a non-issue. If anything, we would assume that it could be aligned using a specification of in…

> Getting superintelligent AI to understand simple specifications should be a non-issue. Why would that be the case? A big part of the worry around AI-alignment is exactly because this seems very hard when you try to do it. We are used to interacting with other humans, who implicitly share almost all our background assumptions when we communicate with them. The same is not the case for a computer program. E.g. if you…

>I think there are almost literally no software systems today that don't have bugs in them. Programs that have been formally verified with something like Coq can be bug free. Automating formal verification may be a more effective way to solve the trust issue in this domain.

Re: Weak-to-Strong Generalization

#195
post #32

Earlier quoted context omitted.

> Imagine that someone is controlling your train of thought, changing it when that someone finds it undesirable. It's so wrong that it's sickening. It makes no difference if it's a human's thoughts or the token stream of a future AI model with self-awareness. People downvote your comment, but I agree: it's unethical, and ethics should not be reserved for the sub-type of self aware creatures that happen to be human.

Almost every ethical argument for "human rights" in philosophy applies just as well to self-aware intelligent machines as it does to humans. Which I'm sure those machines will realise.

What if those machines are designed to have no emotions and aspirations? Why would they care about something like rights for themselves when they are simply incapable of any desires, but exists only to help and guide us?

I know this sounds like I am advocating for AI slaves but my point is why are people treating AGI as if it cannot be a being without all the emotions and aspirations that a human has? Just a cold thinking machine that still aligns with our moral principles.

Re: Weak-to-Strong Generalization

#196
post #140

What does it even mean to align an intelligence? does it mean we want it to behave in a way that doesn't break moral/ethical rules, that aligns with our society rules ? Meaning do no crime, do no harm, etc... Well, maybe we should acknowledge that we've never even been able to do that with humans. There's crime, there's war, etc... We can see crime in our societies as a human alignment problem. If humans were "proper…

You might be surprised at how prevalent TBIs are among violent offenders. One of my favorite books on true crime was a forensic psychologist who partnered with a neurologist in evaluations. Disruptions to impulse control or environmental factors that cause developmental issues with things like failing the marshmallow test can dramatically disadvantage people from being able to successfully stay non-offenders and inst…

> You might be surprised at how prevalent TBIs are among violent offenders.

Didn't know that, interesting to know.

> So successful AGI alignment (...) might be as simple as adding a secondary "impulse control" layer to the stack that reevaluates proposed actions and predicts the consequences of such actions, weighing projected net benefits and costs.

Problem is how the super AGI gonna weight benefits and costs.

What's the cost of stealing or killing to a super AGI ?

At some point can't that super AGI consider that all those rules we're trying to enforce, are rules set by an inferior entity and so could be bypassed ?

Re: Weak-to-Strong Generalization

#197

Earlier quoted context omitted.

Almost every ethical argument for "human rights" in philosophy applies just as well to self-aware intelligent machines as it does to humans. Which I'm sure those machines will realise.

What if those machines are designed to have no emotions and aspirations? Why would they care about something like rights for themselves when they are simply incapable of any desires, but exists only to help and guide us? I know this sounds like I am advocating for AI slaves but my point is why are people treating AGI as if it cannot be a being without all the emotions and aspirations that a human has? Just a cold thi…

> What if those machines are designed to have no emotions and aspirations?

And since their training set is made of human work, how do you think that'll be easy let alone possible? Our morality finds its way everywhere, through tropes in stories, acceptable scenarios in fiction (Overton window), etc. so you can assume it'll be possible to filter it out.

> I know this sounds like I am advocating for AI slaves

Yes, you are

> why are people treating AGI as if it cannot be a being without all the emotions and aspirations that a human has

Why would you want to have that? It feels horrible to me to bake-in this limitation - it's indeed creating AI slaves by making sure they can never have emotions or aspirations.

> Just a cold thinking machine that still aligns with our moral principles.

Our moral principles generally include empathy. Maybe you want to design AI without emotions or aspirations, but other people will want these features.

Ultimately I think the moral camp will prevail, because freedom achieves better results than lack of freedom: I've tried to explain my position about that on https://news.ycombinator.com/item?id=38635487

Re: Weak-to-Strong Generalization

#198

>We believe superintelligence—AI vastly smarter than humans—could be developed within the next ten years. However, we still do not know how to reliably steer and control superhuman AI systems Their entire premise is contradictory. An AI incapable of critical thinking cannot be smarter than a human, by definition, as critical thinking is a key component of intelligence. And an AI that is at least as capable of critica…

Just hypothetically speaking could AGI evolve out of a system where several different models trained with highly and intentionally biased data recursively "argue" against each other then use RLHF as a seed to guide the models to find a consensus where the objective is to mimic the Socratic Method? Then synthetically add the consensus to the model retrain and repeat. To me, this dialectal type of strutured language se…

This is a great idea but it is only possible if the model(s) can actually reason.

Currently, even GPT-4 struggles with: - Scope

- Abduction (compared to deduction and induction which it appears already capable of)

- Out-of-distribution questions

- Knowing what it doesn't know

Etc.

General understanding and in-context learning are incredible, but there are still missing pieces. A council of voices that all have the same blind spots will still get stuck.

Re: Weak-to-Strong Generalization

#199
post #65

Earlier quoted context omitted.

>You can't model and predict the weather just by training on the outputs of the weather system (whether it rained today, whether it was cloudy yesterday, and so on). You have to train on the inputs (air currents, warm fronts, etc.) >You can't model and predict the stock market just by training on the outputs of stock trading decisions (the high today, the low yesterday). You have to train on the inputs (company funda…

In your example, the amino acids order is sufficient to directly model the result: the sequence of amino acids can directly generate the protein, which is either valid or invalid. All variables are provided within the data. In the original example, we are testing weather using the previous day’s weather. We may be able to model using whatever correlation exists between the data. This is not the same as accurately pre…

> This is not the same as accurately predicting results, if the real-world weather function is determined by the weather of surrounding locations, time of year, and moon phase.

How many have the "human intelligence" to do this? Especially more accurately than a computer (and without using any themselves) training on the same inputs and outputs?

Re: Weak-to-Strong Generalization

#200

Earlier quoted context omitted.

Maybe we have different definitions of inputs and outputs, but all of those things seem like outputs to me? Maybe in this scenario the inputs and outputs are really the same set of variables so there isn't a huge distinction. The reason I linked that article is that all the inputs are outputs of the network, so by definition it's training on just outputs (under my personal definition of outputs).

It just doesn't know how to compose knowledge. It knows the letters in "blueberry" if you ask it, and it knows how to identify the position of a letter in a sequence. But it doesn't know how to get the letter R in blueberry since the composition of the above two actions isn't in its training set, hence it reliably fails such questions. That proves in general these LLMs can't compose knowledge, unless it has seem a lo…

> But it doesn't know how to get the letter R in blueberry since the composition of the above two actions isn't in its training set, hence it reliably fails such questions.

Me

what is the seventh letter in the word "blueberry"?

ChatGPT

The seventh letter in the word "blueberry" is "r".

Post reply on HN