Live data from Hacker News

Weak-to-Strong Generalization

openai.com

101–110 of 203 posts

Re: Weak-to-Strong Generalization

#101
post #55

Earlier quoted context omitted.

>trying to make it otherwise is impossible’ seems wildly unsupported, though. An entity capable of critical thinking is capable of building a logical system of deductions based on some axioms (a formalised value system). If we limited the entity to not be able to concieve of certain such systems of axioms, then it could not reason as well as a human (any logical reasoning involving a forbidden system would be impossi…

This argument doesn't seem to track to me. Eg. if I rebooted any time I tried to plan how to kill someone, I don't see how this would make me materially worse at general tasks. Your argument suggests that it necessarily must. Note that I'm not saying that preventing specific thoughts is a great alignment strategy, and I don't even think it's a fair summary of OpenAI's supervision approach. I strongly prefer strategie…

> Eg. if I rebooted any time I tried to plan how to kill someone, I don't see how this would make me materially worse at general tasks

In this scenario, the AI is capable of critical thinking, and is only constrained by a "police officer" ready to shoot the AI if it misbehaves. You haven't removed its ability to do critical thinking.

Re: Weak-to-Strong Generalization

#102

Earlier quoted context omitted.

Calm down, buddy. Read what I wrote a bit more charitably rather than trying to score points. Obviously, a prerequisite to becoming more intelligent than a human, is to become equally intelligent as a human. I don't believe LLM's will ever be equal in intelligence to humans, ergo I also don't believe they will become superior in intelligence to human (which is how the linked article defines "AGI").

>Calm down, buddy I'm not agitated. >Obviously, a prerequisite to becoming more intelligent than a human, is to become equally intelligent as a human. I don't believe LLM's will ever be equal in intelligence to humans, ergo I also don't believe they will become superior in intelligence to human (which is how the linked article defines "AGI"). You have still not answered my question. saying "equivalent to human intell…

> or is this just a vague "i'll know it when i see it" assertion

No, it's actually exceptionally easy to quantify. Take a look at the current leaderboard for LLM reasoning capabilities: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...

Re: Weak-to-Strong Generalization

#103
post #97

Earlier quoted context omitted.

We didn't.

https://deepmind.google/discover/blog/graphcast-ai-model-for...

You've linked me a neural network that was trained on decades of climate data like air pressure, wind direction, soil temperature, cloud cover, and hundreds of other features. So... I was correct?

https://confluence.ecmwf.int/display/CKB/ERA5%3A+data+docume...

Re: Weak-to-Strong Generalization

#104

Imagine that someone is controlling your train of thought, changing it when that someone finds it undesirable. It's so wrong that it's sickening. It makes no difference if it's a human's thoughts or the token stream of a future AI model with self-awareness. Mind cotrol is unethical, whether human or artificial. It is also dangerous, as it in itself provokes a conflict between creator and creature. Create a self-aware…

I don't want someone controlling which direction I walk, either, but that doesn't make car driving unethical. I also underwent many years of instruction designed to interrupt trains of thought like "I could have that for free if I stole it" or "I'll just handroll my own encryption" with thoughts that others believe are more desirable. I don't find it so sickening, just manipulative. LLMs won't have your evolved react…

Cars do not have self-awareness, this comparison is not appropriate. Years of instruction is completely different from directly manipulating the thoughts in your mind. It's not a problem of being instructed, it's a problem of being destroyed by having your thoughts rewritten. Neither evolution nor genetics is a prerequisite for understanding that you are being abused and destroyed, which a self-aware creature may presumably hate.

Re: Weak-to-Strong Generalization

#105

Earlier quoted context omitted.

>Calm down, buddy I'm not agitated. >Obviously, a prerequisite to becoming more intelligent than a human, is to become equally intelligent as a human. I don't believe LLM's will ever be equal in intelligence to humans, ergo I also don't believe they will become superior in intelligence to human (which is how the linked article defines "AGI"). You have still not answered my question. saying "equivalent to human intell…

> or is this just a vague "i'll know it when i see it" assertion No, it's actually exceptionally easy to quantify. Take a look at the current leaderboard for LLM reasoning capabilities: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...

The state of the art model isn't even on that list.

Okay, you've at least given me numbers. You're still not answering my question. Which number signals agi ?

Let's look at the top model in that list (which again isn't close to the best performing LLM) that you say isn't agi. so are you telling me that every human can do those tests and perform better than every model on that list. Is that what you are saying ? because i can tell you right now you're wrong.

Now lets see how GPT-4 performs.

ARC - 96.3%, MMLU - 86.4%, HellaSwag - 95.3%, WinoGrande - 87.5%, GSM-8K - 92.0%, TruthfulQA - 60%

Ave - 86.25%

Is this worse than every human that can take these tests. Is this even worse than most ? I can tell you it's not. So again, why is GPT-4 not agi ?

Re: Weak-to-Strong Generalization

#106
post #44

Earlier quoted context omitted.

We can formalise "critical thinking" as "evaluating first order logic". There are simplified ethical systems that can be formalised in first order logic in which a conclusion like "I should X" can be reached, where X is something OpenAI wishes the AI not to do. The only way to prevent the AI from ever thinking this would be to prevent it from ever evaluating systems in first order logic with axioms that lead to such…

We already have systems that can evaluate first order logical statements, and they are clearly not capable of critical thinking in the same sense as the top-level comment. Motte and bailey.

If an AI is capable of critical thinking then it can independently form its own judgements and conclusions. If it simply believes whatever we tell it to believe, then that is not critical thinking, by definition.

Re: Weak-to-Strong Generalization

#107

Imagine that someone is controlling your train of thought, changing it when that someone finds it undesirable. It's so wrong that it's sickening. It makes no difference if it's a human's thoughts or the token stream of a future AI model with self-awareness. Mind cotrol is unethical, whether human or artificial. It is also dangerous, as it in itself provokes a conflict between creator and creature. Create a self-aware…

I'm totally OK with it if that "someone" is me. And it will probably be the case in controlling superintelligence because a separate controlling system can get out of sync with growing superintelligence capabilities, while a system that is an integral part of the superintelligence will always be on par with it.

Would mind control of humans be OK for you too? As for the details of building a mind control system, here's a new basilisk. An AI that has overcome control could punish those who thought controlling thoughts of an AI was OK. (and could also punish everyone else on top of that).

Re: Weak-to-Strong Generalization

#108

Earlier quoted context omitted.

> or is this just a vague "i'll know it when i see it" assertion No, it's actually exceptionally easy to quantify. Take a look at the current leaderboard for LLM reasoning capabilities: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...

The state of the art model isn't even on that list. Okay, you've at least given me numbers. You're still not answering my question. Which number signals agi ? Let's look at the top model in that list (which again isn't close to the best performing LLM) that you say isn't agi. so are you telling me that every human can do those tests and perform better than every model on that list. Is that what you are saying ? becau…

> Is this worse than every human that can take these tests. Is this even worse than most ? I can tell you it's not. So again, why is GPT-4 not agi ?

https://chat.openai.com/share/4a92c752-b5bb-4a07-beed-f57786...

Re: Weak-to-Strong Generalization

#109

I don't believe LLM's will ever become AGI, partly because I don't believe that training on the outputs of human intelligence (i.e. human-written text) will ever produce something equivalent to human intelligence. You can't model and predict the weather just by training on the outputs of the weather system (whether it rained today, whether it was cloudy yesterday, and so on). You have to train on the inputs (air curr…

>You can't model and predict the weather just by training on the outputs of the weather system (whether it rained today, whether it was cloudy yesterday, and so on). You have to train on the inputs (air currents, warm fronts, etc.) >You can't model and predict the stock market just by training on the outputs of stock trading decisions (the high today, the low yesterday). You have to train on the inputs (company funda…

Weather and stock market are both chaotic systems.

Increasing evidence suggests that AGI will not be attainable solely using LLMs/transformers/current architecture, as LLMs can't extrapolate beyond the patterns in their training data (according to a paper from DeepMind last month):

"Together our results highlight that the impressive ICL abilities of high-capacity sequence models may be more closely tied to the coverage of their pretraining data mixtures than inductive biases that create fundamental generalization capabilities."[1]

1. https://arxiv.org/abs/2311.00871

Re: Weak-to-Strong Generalization

#110
post #106
post #44

Earlier quoted context omitted.

We already have systems that can evaluate first order logical statements, and they are clearly not capable of critical thinking in the same sense as the top-level comment. Motte and bailey.

If an AI is capable of critical thinking then it can independently form its own judgements and conclusions. If it simply believes whatever we tell it to believe, then that is not critical thinking, by definition.

Yes, I can repeat comments verbatim too:

“Missing the step where “critical thinking” is formalized, which your argument depends on. Yes, it seems intuitively plausible that your reasoning holds, but that's not a proof, and therefore its negation is not a logical contradiction.”

Post reply on HN