Eight things to know about large language models [pdf]
1–10 of 114 posts
Re: Eight things to know about large language models [pdf]
#2Its actually a crazy idea that the only thing we needed to do to make LLMs more effective was just scale up the parameters
Though the transformer architecture definitely made it possible to handle the complexity
Re: Eight things to know about large language models [pdf]
#3Leaving that aside, I really take issue with the style used by the author. For example, section 3 begins:
> There is increasingly substantial evidence that LLMs develop internal representations of the world to some extent, and that these representations allow them to reason at a level of abstraction that is not sensitive to the precise linguistic form of the text that they are reasoning about.
LLMs do not "reason"; they do not "learn" or "develop" anything of their own volition. They are (advanced) statistical models. Anthropomorphizing them is not just technically incorrect, but morally disingenuous. Writing about LLMs in this way causes people with less domain-specific knowledge to trust the models the way they might trust people, and this has the potential for serious harm.
Because of this choice of phrasing, I wanted to look into the author's background. Among their recent activities, they list:
> I'm organizing a new AI safety research group at NYU, and I wrote up a blog post explaining what we're up to and why.
"AI Safety" is a distinct (and actually opposing) area from "AI Ethics". The people who prefer the word "safety" tend also to engage in discussions that touch on aspects of longtermism. Longtermism is not scientifically well-grounded; it seeks to divert attention from present and real issues to fanciful projections of far-future concerns. I do not know for certain that the author is in fact a longtermist, but their consistent anthropomorphization of a pile of statistical formulae certainly suggests they wouldn't feel out of place among a crowd of such people.
In contrast, the people who prefer the term "ethics" in their work are grounded in real and present issues. They concern themselves with reasonable regulation. They worry about the current-day environmental impacts of training large models. In short, they are concerned with actual issues, rather than the alleged potential for a statistical model to "develop" sentience or exhibit properties of "emergent" intelligence (subjects from the annals of the science-fiction writing of last century).
I hope the author can clarify their choice of phrasing in their work, though I worry they have chosen their words carefully already. Readers should exercise caution in taking the claims of a soothsayer without a sufficient quantity of salt.
Re: Eight things to know about large language models [pdf]
#4This is a personal correspondence typeset via LaTeX — it is not an academic paper, and it was not peer-reviewed. (The document does not claim otherwise, but I think it's common for people to assume that documents that have been typeset in such a format are more rigorous than this is.) Leaving that aside, I really take issue with the style used by the author. For example, section 3 begins: > There is increasingly subs…
Re: Eight things to know about large language models [pdf]
#5On the one hand, Gaining capabilities unpredictably is a bit exaggerated - it's more that a lot of apparent capabilities are embedded in language and these models approximate the truly vast amounts of text they've digested (imo). Just as much, LLMs don't express their creators values 'cause they don't express any values, they average language responses (though people certainly can see them expressing values, which can cause problems for the people).
On the other hand, the point that their creators don't understand them and are still quite willing to toss them to the open Internet really show considered carefully whatever their exact capabilities.
Re: Eight things to know about large language models [pdf]
#6Would that enabled these models to compress themselves?
Re: Eight things to know about large language models [pdf]
#7This is a personal correspondence typeset via LaTeX — it is not an academic paper, and it was not peer-reviewed. (The document does not claim otherwise, but I think it's common for people to assume that documents that have been typeset in such a format are more rigorous than this is.) Leaving that aside, I really take issue with the style used by the author. For example, section 3 begins: > There is increasingly subs…
What definition of reasoning do you have in mind such that LLMs don't do it?
Re: Eight things to know about large language models [pdf]
#8This is a personal correspondence typeset via LaTeX — it is not an academic paper, and it was not peer-reviewed. (The document does not claim otherwise, but I think it's common for people to assume that documents that have been typeset in such a format are more rigorous than this is.) Leaving that aside, I really take issue with the style used by the author. For example, section 3 begins: > There is increasingly subs…
What does “reason” mean? It seems like it does everything I expect from something that reasons.
Seriously just watch. He's not actually going to be able to coherently define his "reasoning" in a way that can be tested.
Re: Eight things to know about large language models [pdf]
#9This is a personal correspondence typeset via LaTeX — it is not an academic paper, and it was not peer-reviewed. (The document does not claim otherwise, but I think it's common for people to assume that documents that have been typeset in such a format are more rigorous than this is.) Leaving that aside, I really take issue with the style used by the author. For example, section 3 begins: > There is increasingly subs…
What does “reason” mean? It seems like it does everything I expect from something that reasons.
The alternative is having 1000 different results for different kinds of arrows, and averaging out the results for the ones similar to the input arrow.
An LLM is working on text tokens, it’s trying to give the most statistically common next token based on everything that’s been fed into it. Does that statistical model abstract the objects and concepts it talks about? Eh? I don’t know
Re: Eight things to know about large language models [pdf]
#10While I don't think these claims are entirely correct, I think they are worth considering. On the one hand, Gaining capabilities unpredictably is a bit exaggerated - it's more that a lot of apparent capabilities are embedded in language and these models approximate the truly vast amounts of text they've digested (imo). Just as much, LLMs don't express their creators values 'cause they don't express any values, they a…
The fact is that they acquired abilities no one expected them to acquire. In hindsight you can say it was embedded in language and maybe you could have seen it coming, but it is an empirical fact that this was unexpected and unpredictable beforehand.