Live data from Hacker News

Eight things to know about large language models [pdf]

cims.nyu.edu

101–110 of 114 posts

Re: Eight things to know about large language models [pdf]

#101
post #33

Earlier quoted context omitted.

> LLMs clearly develop internal representations, this is an empirical fact. you've offered an anecdote. Some LLMs (more generally, this type of neural network model) will generate configurations that can be understood as a representation; others will not. The fact that the authors were able to find an apparent model in a heavily rule-based system is not incredibly surprising but offers little clue about whether this…

Okay, I think we are in agreement. Let's say current LLM architectures are clearly capable of developing internal representations, not just learning surface statistics, and it can execute algorithmic computation as complex as computing Othello board states, and it can completely generalize out of distribution thanks to such algorithmic computation. (One experiment was to completely eliminate any Othello games startin…

The othello example explicitly was trained on the rules of the game - it isn't suprising that you can find a representation of game state.

LLMs are trained on the "rules" of human written language, which are sufficiently distinct from anything in the actual world that find world modelling would, indeed, be a surprise to me.

Re: Eight things to know about large language models [pdf]

#102

This is a personal correspondence typeset via LaTeX — it is not an academic paper, and it was not peer-reviewed. (The document does not claim otherwise, but I think it's common for people to assume that documents that have been typeset in such a format are more rigorous than this is.) Leaving that aside, I really take issue with the style used by the author. For example, section 3 begins: > There is increasingly subs…

It's been a long time since anyone reminded me of Richard Bach's 'illusions' Thank you for a pleasant break! https://www.patheos.com/blogs/wakeupcall/2020/05/10-thought-...

It's always a happy moment when someone recognizes my online namesake. Cheers!

Re: Eight things to know about large language models [pdf]

#103
post #55

This is a personal correspondence typeset via LaTeX — it is not an academic paper, and it was not peer-reviewed. (The document does not claim otherwise, but I think it's common for people to assume that documents that have been typeset in such a format are more rigorous than this is.) Leaving that aside, I really take issue with the style used by the author. For example, section 3 begins: > There is increasingly subs…

>LLMs do not "reason"; they do not "learn" or "develop" anything of their own volition. They are (advanced) statistical models. As a non-expert in the field I was hesitant at the time to disagree with the legions of experts who last year denounced Blake Lemoine and his claims. I know enough to know, though, of the AI effect https://en.wikipedia.org/wiki/AI_effect >, a longstanding tradition/bad habit of advances bein…

I didn't say it's "not real AI", but of course whether I meant that or not comes down to a definition of "AI".

In academia (where I am, though my specialty is a step removed), "AI" is exactly the region of research concerned with statistical problem-solving. Machine learning is sometimes synonymous, though I haven't been able to determine whether this is consistent among all self-described AI researchers or just some of them.

Systems like GPTs are not capable of "thought", full stop. There is no ambiguity. If this is hard for you to accept, I'm sorry, but it is a fact. They have no agency.

I saw someone sketch a good explanation. They were responding to a claim that Chat GPT could "pass the Harvard admissions process". While it is true that the system can access sufficient data to respond accurately to questions on an exam, and it could be prompted to generate an essay, it cannot of its own volition choose to submit an application. It doesn't even "know" that Harvard is a real place to which it can apply, nor does it "understand" how it could go about "learning" how to find out about these things it does not know. It simply isn't capable of these things. (And even if it were that wouldn't necessarily be evidence of intelligence, but for now it's a sufficient distinguisher.)

> Wat

Can you explain what about my last point was "wat"-worthy? Do you disagree that it is morally problematic to pretend a thing is sentient when it is in fact not?

Re: Eight things to know about large language models [pdf]

#104
post #91

Earlier quoted context omitted.

It's a word prediction algorithm. Literally any collection of words (sometimes also known as a fact) it was trained on, regardless of how esoteric or domain specific, should generally be able to be regurgitated and, to a lesser degree, associated with similar collections of words. If you want to see it fail, don't try to "outsmart" it, but simply consider how it's programmed. --- Me: "continue the sequence: 0, done,…

ChatGPT is not good when it needs to answer directly. If you let it lead itself to the right answer it's pretty good. # Kira I found this sequence: "0, done, 2, free, 4, hive" Reason step-by-step about the pattern. Consecutively continue the sequence with four more entries. Consecutively reflect on your answer and fix mistakes in case you broke the pattern. # ChatGPT4 After examining the given sequence, "0, done, 2,…

This is a great example of why some people are being paid $350,000/year to prompt LLMs while the rest of us are going to be living in cardboard boxes.

Re: Eight things to know about large language models [pdf]

#105

Earlier quoted context omitted.

… and then perform a careful search of books and the whole internet to be sure what you think is novel hasn’t been thoroughly debated somewhere on stackexchange.

If it's a stochastic parrot, then merely randomizing proper nouns and filler text should be enough to prevent its abstraction ability. If you're saying that we can't use a problem if any analog of that problem has ever been described, you seem to be arguing more strongly that it is a general intelligence than I am.

I’m saying that people in my circle have been asking what they think are novel questions and getting interesting answers, only to find out that very similar content exists on websites we know are in the training set.

That’s not intelligence that’s computers having better memory than humans. Useful, certainly, but hardly skynet.

Re: Eight things to know about large language models [pdf]

#106

Earlier quoted context omitted.

If it's a stochastic parrot, then merely randomizing proper nouns and filler text should be enough to prevent its abstraction ability. If you're saying that we can't use a problem if any analog of that problem has ever been described, you seem to be arguing more strongly that it is a general intelligence than I am.

I’m saying that people in my circle have been asking what they think are novel questions and getting interesting answers, only to find out that very similar content exists on websites we know are in the training set. That’s not intelligence that’s computers having better memory than humans. Useful, certainly, but hardly skynet.

I don't think you're being clear about whether the questions were novel. If you discover your question was uncreative, surely e.g. some details, wording, facts, names, or numbers inside the question can be changed to defeat a model that is answering it from memory?

If you're saying that it is not possible to change the details enough to avoid the model being able to answer that type of question, I think you are admitting that the model has learned a generalized ability to answer questions of that class, and is not actually using its memory to answer at all.

I don't care about whether it learned that generalized ability from seeing examples of the question and answer, which it then deduced an algorithm for and generalized -- that's how most people learn most things.

Re: Eight things to know about large language models [pdf]

#107

Earlier quoted context omitted.

I’m saying that people in my circle have been asking what they think are novel questions and getting interesting answers, only to find out that very similar content exists on websites we know are in the training set. That’s not intelligence that’s computers having better memory than humans. Useful, certainly, but hardly skynet.

I don't think you're being clear about whether the questions were novel. If you discover your question was uncreative, surely e.g. some details, wording, facts, names, or numbers inside the question can be changed to defeat a model that is answering it from memory? If you're saying that it is not possible to change the details enough to avoid the model being able to answer that type of question, I think you are admit…

[deleted]

Re: Eight things to know about large language models [pdf]

#108

Earlier quoted context omitted.

I’m saying that people in my circle have been asking what they think are novel questions and getting interesting answers, only to find out that very similar content exists on websites we know are in the training set. That’s not intelligence that’s computers having better memory than humans. Useful, certainly, but hardly skynet.

I don't think you're being clear about whether the questions were novel. If you discover your question was uncreative, surely e.g. some details, wording, facts, names, or numbers inside the question can be changed to defeat a model that is answering it from memory? If you're saying that it is not possible to change the details enough to avoid the model being able to answer that type of question, I think you are admit…

The asker thought they were. They were not. The internet is big and human memories are not.

As an aside, I’m really starting to hate these threads on here, people are constantly reading words that aren’t there in search of gotcha-it’s-skynet. It’s not. It’s just pattern matching and randomness with a giant amount of information encoded.

Re: Eight things to know about large language models [pdf]

#109
post #95

Earlier quoted context omitted.

You’re not alone - I had to read it out loud to see the pattern :)

I wonder if people have a range of sensitivity to what one might call the "mind's voice" - by analogy to the recent discovery of aphantasia, which revealed that different people's experience of the "mind's eye" ranges from vivid imagery to nothing at all.

Some people say they don't have a mind voice, and grow up thinking nobody does. Other people, with a mind voice, grow up thinking everyone does.

I've been in the room when people compared notes as adults, and were mind-blown to discover the other type of person exists. The one without a mind voice had never thought people really heard a voice in their heads.

Re: Eight things to know about large language models [pdf]

#110

Earlier quoted context omitted.

I don't think you're being clear about whether the questions were novel. If you discover your question was uncreative, surely e.g. some details, wording, facts, names, or numbers inside the question can be changed to defeat a model that is answering it from memory? If you're saying that it is not possible to change the details enough to avoid the model being able to answer that type of question, I think you are admit…

The asker thought they were. They were not. The internet is big and human memories are not. As an aside, I’m really starting to hate these threads on here, people are constantly reading words that aren’t there in search of gotcha-it’s-skynet. It’s not. It’s just pattern matching and randomness with a giant amount of information encoded.

> As an aside, I’m really starting to hate these threads on here.

I'm not sure what to say, other than that if you'd like to have less frustrating conversations, you could do better than showing up with hearsay where someone asked a question they thought was unique, but it wasn't, and it can't be modified to be unique and then asked again, and you aren't willing to tell us what it was, and possibly don't know yourself.

It is not possible to have a serious conversation about your claim, and that's not because it is being intentionally misunderstood.

> skynet

You're the only person mentioning skynet. The conversation is about a ridiculous claim made up-thread that GPT-4 cannot reason or understand anything, which is disprovable within a few minutes of using it thoughtfully.

Post reply on HN