Earlier quoted context omitted.
Emergent properties are never "real". They just are and you can see them happening, but "underneath" it's nothing. Edit: I meant to say I don't need access to training. By experimenting with in/outputs you can get a basic picture. I don't need to see biological scans to say something about your personality either.
I think an important distinction here is to say that currently, you perceive them to be real. They aren't factually real things, at least not yet. Judging someones personality is a subjective process, not an objective one.
Saying Goodbye to GitHub
411–420 of 450 posts
Re: Saying Goodbye to GitHub
#412Earlier quoted context omitted.
On the books? I mean laws like "you shouldn't make money from OSS made by someone else". The context of this chat.
oh, so those things are what people call "opinions", which they are completely allowed to have, just as you have yours. They aren't oppressing you, they don't expose you to penalties, you can't get thrown in jail.
Re: Saying Goodbye to GitHub
#413Earlier quoted context omitted.
> I think we’re now way past that now with LLMs now quickly taking on the role of a general reasoning engine. No we're not, and no they are not. An LLM doesn't reason, period. It mimics reasoning ability by stochastically chosing a sequence of tokens. Alot of the time these make sense. At other times, they don't make any sense. I recently asked an LLM: "Mike leaves the elevator at the 2nd floor. Jenny leaves at the 9…
> but the fact that they can also spew such complete illogical nonsense shows that they are not "reasoning" about things Have you ever seen the proof that 2=1 ? It looks convincing, but it's illogical because it has a subtle flaw. Are the people who can't spot the flaw just " looking like they are reasoning", but really they just lack the ability to reason? Are witnesses who unintentionally make up memories in court…
Lacking relevant information or insight into a topic, isn't the same as lacking the ability to reason.
> You can't just spout that an LLM lacks reasoning without first strictly defining what it means to reason.
Perfectly worded definition available on Wikipedia:
Reason is the capacity of consciously applying logic by drawing conclusions from new or existing information, with the aim of seeking the truth.
"Consciously", "logic", and "seeking the truth" are the operative terms here. A sequence predictor does none of that. Looking at my above example: The sequence "Mike leaves the elevator first" isn't based on logical thought, or a conscious abstraction of the world built from ingesting the question. It's based on the fact that this sequence has statistically a higher chance to appear after the sequence representing the question.How does our reasoning work? How do humans answer such a question? By building an abstract representation of the world based on the meaning of the words in the question. We can imagine Mike and Jenny in that Elevantor, we can imagine the elevator moving, floor numbers have meaning in the environment, and we understand what "something is higher up" means. From all this we build a model and draw conclusions.
How does the "reasoning" in the LLM work? It checks which tokens are likely to appear after another sequence of tokens. It does so by having learned how we like to build sequences of tokens in our language. That's it. There is no modeling of the situation going on, just stochastic analysis of a sequence.
Consequently, an LLM cannot "seek truth" either. If a sequence has a high chance of appearing in a position, it doesn't matter if it is factually true or not, or even logically sound. The model isn't trained on "true or false". It will, likely more often than not say things that are true, but not because it understands truth, but because the training data contain a lot of token sequences that, when interpreted by a human mind, state true things.
Lastly, imagine trying to apply a language model to an area that depends completely on the above definition of reasoning as a consequence of modeling the world based on observations and drawing new conclusions from that modeling.
https://www.spiceworks.com/tech/artificial-intelligence/news...
Re: Saying Goodbye to GitHub
#414Earlier quoted context omitted.
> but the fact that they can also spew such complete illogical nonsense shows that they are not "reasoning" about things Have you ever seen the proof that 2=1 ? It looks convincing, but it's illogical because it has a subtle flaw. Are the people who can't spot the flaw just " looking like they are reasoning", but really they just lack the ability to reason? Are witnesses who unintentionally make up memories in court…
> Are the people who can't spot the flaw just "looking like they are reasoning", but really they just lack the ability to reason? Lacking relevant information or insight into a topic, isn't the same as lacking the ability to reason. > You can't just spout that an LLM lacks reasoning without first strictly defining what it means to reason. Perfectly worded definition available on Wikipedia: Reason is the capacity of c…
The sequence "Mike leaves the elevator first" has a high statistical probability. The sequence "Jenny leaves the elevator first" has a lower probability that that. But it probably has still a much higher probability than "Michael is standing on the Moon", which in turn may be more likely than "Car dogfood sunshine Javascript", which is still probably more likely than "snglub dugzuvutz gummmbr ha tcha ding dong".
Note that none of these sequences are wrong in the world of a language model. They are just increasingly unlikely to occur in that position. To us with our ability to reason by logically drawing conclusions from an abstract internal model of the world, all these other sequences either represent false statements, or nonsensical word sald.
Re: Saying Goodbye to GitHub
#415Earlier quoted context omitted.
There's a class of statements that can be either interpreted precisely, at which point the claim they make is clearly true but trivial, or interpreted expansively, at which point the claim is significant but no longer clearly true. This is one of those: yes, technically LLMs are token predictors, but technically any nondeterministic Turing machine is a token predictor. The human brain could be viewed as a token predi…
> The human brain could be viewed as a token predictor No it really couldn't, because "generating and updating a 'mental model' of the environment." is as different from predicting the next token in a sequence, as a bees dance is from a structured human language. The mental model we build and update is not just based on a linear stream, but many parallel and even contradictory sensory inputs that we make sense of not…
> The mental model we build and update is not just based on a linear stream, but many parallel and even contradictory sensory inputs
So just like multimodal language models, for instance GPT-4?
> as experiences in a world of which we are part of.
> The simple fact that we don't just complete streams, but do so with goals, both immediate and long term, and fit our actions into these goals
Unfalsifiable! GPT-4 can talk about its experiences all day long. What's more, GPT-4 can act agentic if prompted correctly. [2] How do you qualify a "real goal"?
[1] https://www.neelnanda.io/mechanistic-interpretability/othell...
Re: Saying Goodbye to GitHub
#416Earlier quoted context omitted.
> The human brain could be viewed as a token predictor No it really couldn't, because "generating and updating a 'mental model' of the environment." is as different from predicting the next token in a sequence, as a bees dance is from a structured human language. The mental model we build and update is not just based on a linear stream, but many parallel and even contradictory sensory inputs that we make sense of not…
But the human mental model is purely internal. For that matter, there is strong evidence that LLMs generate mental models internally. [1] Our interface to motor actions is not dissimilar to a token predictor. > The mental model we build and update is not just based on a linear stream, but many parallel and even contradictory sensory inputs So just like multimodal language models, for instance GPT-4? > as experiences…
Limited models, such as those representing the state of a game that it was trained to do: Yes. This is how we hope deep learning systems work in general.
But I am not talking about limited models. I am talking about ad-hoc models, built from ingesting the context and semantic meaning of a string of tokens, that can simulate reality and allows drawing logical conclusions from it.
In regard to my example given elsewhere in this HN thread: I know that Mike exits the elevator first because I build a mental model of what the tokens in the question represent. I can draw conclusions from that model, including new conclusions whos token-representation would be unlikely in the LLMs model, which doesn't explain anything about reality, but explains how tokens are usually ordered in the training set.
Re: Saying Goodbye to GitHub
#417Earlier quoted context omitted.
> But being better at mimicking reason, is still not reasoning How do I know people are not using a similar process when they perform "reasoning" but with a way more elaborate model? Can you prove me that the two are inherently different in the type of output they produce regardless of how large a ML model is or can be? Because if you can't, and they produce the same type of output, the processing could be similar en…
> but with a way more elaborate model? Simple: I know that humans have intentionality and agency. They want things, they have goals both immediate and long term. Their replies are based not just on the context of their experiences and the conversation but their emotional and physical state, and the applicability of their reply to their goals. And they are capable of coming up with reasoning about topics for which the…
This all seems orthogonal to reasoning, but also who is to say that somewhere in those billions of parameters there isn't something like a model of goals and emotional state? I mean, I seriously doubt it, but I also don't think I could evidence that.
Re: Saying Goodbye to GitHub
#418Earlier quoted context omitted.
I agree about the utility part. However, I don't really accept the idea that this isn't reasoning, but I'm not entirely sold either way. I'd say if it mimics something well enough then eventually it's just doing the thing, which is the same side of the argument I fall on with Searle's Chinese Room Argument. If you can't discern a difference, is there a difference? So far GPT-4 can produce better work than like 50% of…
> I'd say if it mimics something well enough then eventually it's just doing the thing Right up to the point where it actually needs to reason, and the mimickry doesn't suffice. My above example about the Football and the Coffemug is an easy one, the objects are well represented in its training data. What if I need a reason why the Service Ping spikes every 60 seconds, here is the code, please LLM look it up. I am su…
Re: Saying Goodbye to GitHub
#419I wonder if there's a licence out there already that enables use for humans, but restricts use by robots. (GPLv4 perhaps?)
Re: Saying Goodbye to GitHub
#420Earlier quoted context omitted.
> All this Free Software movement started by something really similar to "right to repair", a firmware bug in a printer that was proprietary software. Free Software is about being in control of software you use. The spirit was never "contribute back to GNU", the spirit was always "if you take GNU software, you can't make it non-free". Those GNU devs at the time just wanted a good and actually free/libre OS, that woul…
It is a pretty big distinction with different end results in practice. Look at Android, you can use the source in a "right to repair" manner but Google doesn't take patches so you can't give back even if you wanted to. The same goes for Apple and Google's OSS browser. The source is there, but there is more or less no way to give back, and they certainly don't.
Well, license obliging the original author to take patches back would be weird one.
But Google could suck in any change to their own and make it better.
> The same goes for Apple and Google's OSS browser. The source is there, but there is more or less no way to give back, and they certainly don't.
That's a different problem that's a bit orthogonal to licensing and has more to do with project leadership. Like, you don't even need to have OSS license to allow users to contribute to project.