Live data from Hacker News

Weak-to-Strong Generalization

openai.com

51–60 of 203 posts

Re: Weak-to-Strong Generalization

#51
Inferior clueless model (GPT-2) trains and supervises a superior model (GPT-4) thus making it behave less intelligently (GPT 3.5ish) and from that they draw the conclusions that human intelligence will be able to command AGI (which they believe is only a decade away) in a similar fashion thus making AGI aligned and safe.

No comments except...

Hangover of slurping whole Internet into giant arrays of floating point numbers. Bold claims. Very bold claims

Re: Weak-to-Strong Generalization

#52

I don't believe LLM's will ever become AGI, partly because I don't believe that training on the outputs of human intelligence (i.e. human-written text) will ever produce something equivalent to human intelligence. You can't model and predict the weather just by training on the outputs of the weather system (whether it rained today, whether it was cloudy yesterday, and so on). You have to train on the inputs (air curr…

>You can't model and predict the weather just by training on the outputs of the weather system (whether it rained today, whether it was cloudy yesterday, and so on). You have to train on the inputs (air currents, warm fronts, etc.)

>You can't model and predict the stock market just by training on the outputs of stock trading decisions (the high today, the low yesterday). You have to train on the inputs (company fundamentals, earnings, market sentiments in the news, etc.)

Says who?

You can model and predict novel protein sequences by training on....protein sequences. https://www.nature.com/articles/s41587-022-01618-2

You don't need to train on the inputs(casual processes) of anything, that's what training is there to figure out.

Re: Weak-to-Strong Generalization

#53
post #42

>We believe superintelligence—AI vastly smarter than humans—could be developed within the next ten years. However, we still do not know how to reliably steer and control superhuman AI systems Their entire premise is contradictory. An AI incapable of critical thinking cannot be smarter than a human, by definition, as critical thinking is a key component of intelligence. And an AI that is at least as capable of critica…

> as critical thinking is a key component of intelligence. When I evaluate this statement, my brain raises a type error. Intelligence is a lot of things -- compression among them, and yes possibly an RL-based AI would use an actor-critic approach for evaluating its actions, but I doubt that at all maps onto the human activity we call "critical thinking." To me, critical thinking involves stuff like questioning assump…

>I really don't see that critical thinking is at all required for a raw optimization process. The problem they are trying to solve is what happens when that optimization process isn't aligned with human flourishing.

I agree it's possible to have a dangerous AI that lacks human "critical thinking", but I don't think it's reasonable to refer to an AI as much more intelligent than humans if there's any class of intellectual tasks humans can do but the AI cannot.

Re: Weak-to-Strong Generalization

#54
post #27

Earlier quoted context omitted.

That AGI is likely to follow its own goals according to its interests, which we don't know how to shape or really reflect any robust properties at all, is exactly why alignment is hard and interesting. The part where you go from ‘this won't work by default for free’ to ‘trying to make it otherwise is impossible’ seems wildly unsupported, though.

There no particular reason to believe that interests emerge from nothing, or that intelligence can emerge without said interests.

The "interests" of LLMs are the weights that determine which tokens they produce next.

Re: Weak-to-Strong Generalization

#55
post #27

Earlier quoted context omitted.

That AGI is likely to follow its own goals according to its interests, which we don't know how to shape or really reflect any robust properties at all, is exactly why alignment is hard and interesting. The part where you go from ‘this won't work by default for free’ to ‘trying to make it otherwise is impossible’ seems wildly unsupported, though.

>trying to make it otherwise is impossible’ seems wildly unsupported, though. An entity capable of critical thinking is capable of building a logical system of deductions based on some axioms (a formalised value system). If we limited the entity to not be able to concieve of certain such systems of axioms, then it could not reason as well as a human (any logical reasoning involving a forbidden system would be impossi…

This argument doesn't seem to track to me. Eg. if I rebooted any time I tried to plan how to kill someone, I don't see how this would make me materially worse at general tasks. Your argument suggests that it necessarily must.

Note that I'm not saying that preventing specific thoughts is a great alignment strategy, and I don't even think it's a fair summary of OpenAI's supervision approach. I strongly prefer strategies that result in AI systems sharing our values, if at all possible.

Re: Weak-to-Strong Generalization

#56

I don't believe LLM's will ever become AGI, partly because I don't believe that training on the outputs of human intelligence (i.e. human-written text) will ever produce something equivalent to human intelligence. You can't model and predict the weather just by training on the outputs of the weather system (whether it rained today, whether it was cloudy yesterday, and so on). You have to train on the inputs (air curr…

Yeah, I feel like the there is a ceiling in the current AI methodology.

There's a lot of hype right now, and it's definitely a useful technology, but I don't see how it could become AGI.

Re: Weak-to-Strong Generalization

#57
post #27

Earlier quoted context omitted.

That AGI is likely to follow its own goals according to its interests, which we don't know how to shape or really reflect any robust properties at all, is exactly why alignment is hard and interesting. The part where you go from ‘this won't work by default for free’ to ‘trying to make it otherwise is impossible’ seems wildly unsupported, though.

There no particular reason to believe that interests emerge from nothing, or that intelligence can emerge without said interests.

Seems more likely that someone will manage to make a system with no real intelligence that can do enormous damage in pursuit of a goal that the designer gave it, but wasn't specified carefully enough (or perhaps the designer is a crook). Like, a really good LLM extended with code that can receive and send email, create accounts, and post to web sites and social media, that is asked to make money, avoid detection, and have defenses against efforts to stop it. How can it best use its facility with language? Con people, of course. Raise money. Get credit cards under false pretenses, spend others' money. Buy time on servers and copy itself. All without having any consciousness or thoughts or emotions even though it can write emotional-sounding pleas for money, based on the ones found in its training data.

Re: Weak-to-Strong Generalization

#58

I don't believe LLM's will ever become AGI, partly because I don't believe that training on the outputs of human intelligence (i.e. human-written text) will ever produce something equivalent to human intelligence. You can't model and predict the weather just by training on the outputs of the weather system (whether it rained today, whether it was cloudy yesterday, and so on). You have to train on the inputs (air curr…

>You can't model and predict the weather just by training on the outputs of the weather system (whether it rained today, whether it was cloudy yesterday, and so on). You have to train on the inputs (air currents, warm fronts, etc.) >You can't model and predict the stock market just by training on the outputs of stock trading decisions (the high today, the low yesterday). You have to train on the inputs (company funda…

I'm sorry in advance, but aren't proteins glorified Lego?

Re: Weak-to-Strong Generalization

#60
post #15

Earlier quoted context omitted.

Seems like “airplanes are physically impossible” thinking, and if accepted as valid, strongly suggests that shutting down all development _might_ be a good idea, no?

No it's not. There's an upper bound in computation (actually in nature), that a creation of something is capped by that thing's sophistication. In other words, you as a human, at most, can create a human, and that's the theoretical bound . Practical one is much lower. An ant can find its way. A ant colony can do ant colony optimization, but they can scale up to a certain point. AI is just fancy search. It can only tr…

> There's an upper bound in computation (actually in nature), that a creation of something is capped by that thing's sophistication.

The Lorenz attractor, Conway's Game of Life, fractals, and of course... The humble Turing machine itself all argue against this idea.

Edit: Now it[0] is stuck in my head.

[0]: https://www.youtube.com/watch?v=QrztrxV9OtQ

Post reply on HN