The number of people ITT who are reading interpreting ChatGPTs output as intelligence is too damn high. I thought Ex Machina was unrealistic because of its dependence on AGI, or at least having a theory of mind. As it turns out, in the real world,a LLM trained on Tinder data could probably get the job done.
What is ChatGPT doing and why does it work?
351–360 of 518 posts
Re: What is ChatGPT doing and why does it work?
#352Earlier quoted context omitted.
That Kenneth Li Othello paper is great. The accompanying blog post https://thegradient.pub/othello/ was discussed on HN here https://news.ycombinator.com/item?id=34474043 A lot of people didn't seem to get it when it was discussed on HN. A GPT had _only_ ever seen Othello transripts like: "E3, D3, C4 ..." and NOTHING else. It knows nothing of the board. It doesnt event know that there are two players. It learned Othe…
I agree this is an incredibly interesting paper. I am not a practitioner but I interpreted the gradient article differently. They didn’t directly find 64 nodes (activations) that represented the board state as I think you imply. They trained “64 independent two-layer MLP classifiers to classify each of the 64 tiles”. I interpret this to mean all activations are fed into a 2 layer MLP with the goal of predicting a sin…
It turns out that the error rates of these probes are reduced from 26.2% on a randomly-initialized Othello-GPT to only 1.7% on a trained Othello-GPT. This suggests that there exists a world model in the internal representation of a trained Othello-GPT.
I take that to mean that the 64 trained Probes are then shown other OthelloGTP internals and can tell us what what the state of their particular 'square' is 98.3% of the time. (we know what the board would look like, but the probes dont)
As you say "Again, not a practitioner but once you are indirecting internal state through a 2 layer MLP it gets less obvious to me that the world model is really there."
But then they go back and actually mess around with OthelloGTPs internal state (using the Probes to work out how), changing black counters to white and so on, and then this directly affects the next move OthelloGTP makes. They even do this for impossible board states (e.g. two unlinked sets of discs) and OthelloGTP still comes up with correct next moves.
So surely this proves that the Probes were actually pointing to an internal model? Because when you mess with the model in a way to affect the next move, it changes OthelloGTPs behaviour in the expected way?
Re: What is ChatGPT doing and why does it work?
#353Earlier quoted context omitted.
That Kenneth Li Othello paper is great. The accompanying blog post https://thegradient.pub/othello/ was discussed on HN here https://news.ycombinator.com/item?id=34474043 A lot of people didn't seem to get it when it was discussed on HN. A GPT had _only_ ever seen Othello transripts like: "E3, D3, C4 ..." and NOTHING else. It knows nothing of the board. It doesnt event know that there are two players. It learned Othe…
> Inside its 'mind', by looking for correlations between its internal state and what they knew the 'board' would look like at each step in the games, they found 64 nodes that seemed to represent the 8x8 Othello board and representation of the two different colours of counters. Is that really surprising though? Take a bunch of sand, and throw it on an architectural relief, and through seemingly random process for each…
Re: What is ChatGPT doing and why does it work?
#354The easiest way for ChatGPT to generate good output is to plain understand it. Given the vast amount of input data fed into it, it has no choice, but to start reducing the input into fundamental rules which is basically what understanding is. Understanding is a form of compression. More efficient for a neural network to understand a concept than memorize permutations. Same with statistics and markov chains, people fo…
Totally agree. I think Juergen Schmidthuber has developed a lot of ideas around compression being the basis for consciousness and understanding. There was the paper that showed that when showing a language model Othello moves it ends up building an internal representation of the board. And now I was reading this abstract: ```Theory of mind (ToM), or the ability to impute unobservable mental states to others, is centr…
Re: What is ChatGPT doing and why does it work?
#355The answer to this is: "we don't really know as its a very complex function automatically discovered by means of slow gradient descent, and we're still finding out" Here are some of the fun things we've found out so far: - GPT style language models try to build a model of the world: https://arxiv.org/abs/2210.13382 - GPT style language models end up internally implementing a mini "neural network training algorithm" (…
Such a hacker news comment. The title isn't a question Stephen Wolfram is asking you a question, it's the title of an article he's written that answers the question.
An ideal hacker news comment would be the exact opposite, it would refer to the article.
Re: What is ChatGPT doing and why does it work?
#356Earlier quoted context omitted.
I can print to PDF as well, but that captures the entire page and I really only want the article. For example, I don’t want the “Recent Writings” column to the right of the article.
The print version has the Recent Writings rendered at the end of the document. On my machine with "US Letter" page size and "No header/footer" I got 72 pages. The main article was 69 pages and the last three were the "cruft" (which includes the Recent Writings). A quick hack might be to postprocess the PDF (eg: using Ghostscript) and trim the last three pages, if all you want is the main article.
I did get what I want with the OneNote clipper.
Edit: Now when I try to print this article to PDF from my phone or iPad, the browser crashes immediately. Something weird is going on with my devices...
Re: What is ChatGPT doing and why does it work?
#357To me, modern AI is just "black boxes all the way down". Even specialists don't really know what's happening. It's not encouraging or interesting. Personally, I'm more interested in analyzing those black boxes than tinkering ones that "seems to work", would it be with graph theory, analysis, etc. To me, if something works but we're unable to really understand why it does, it's more the realm of "testing broken clocks…
I agree. I’m currently taking a neural networks course and it’s making my discomfort worse, not alleviating it. At the beginning of the course we talked about biological models of neurons and that was pretty cool, if a bit simplistic. Now we’re deep into automatic differentiation and gradient descent and a bunch of hidden layers. Ultimately it’s all just using calculus to approximate some unknown function given a sam…
Automatic differentiation is too deep in the weeds whereas the 'behaviour' of neural networks is more emergent.
Re: What is ChatGPT doing and why does it work?
#358The answer to this is: "we don't really know as its a very complex function automatically discovered by means of slow gradient descent, and we're still finding out" Here are some of the fun things we've found out so far: - GPT style language models try to build a model of the world: https://arxiv.org/abs/2210.13382 - GPT style language models end up internally implementing a mini "neural network training algorithm" (…
I will start the first word in French, the second word in English, the third and fourth one in Portuguese, then Spanish, ending up with an Italian verb and German while concluding with a Dutch word. All this while trying to build a grammatically correct question. A bit of a stretch but can be made to work.
The quality of the model answers, does not seem to suffer. It's interesting to see how adding different languages in different point of the phrased question will trigger it to start answering on a different language.
Re: What is ChatGPT doing and why does it work?
#359Earlier quoted context omitted.
It's clearly early technology, so it's not perfect. But what it is able to get right is clear proof it's more then what you think: https://www.engraved.blog/building-a-virtual-machine-inside/ Read to the end. The beginning and middle doesn't show off anything too impressive. It's the very end where chatGPT displays a sort of self awareness. Also here's a scientific paper showing that LLMs are more then a chinese room…
>... chatGPT displays a sort of self awareness. If you read books or articles you will find many places where it appears that whoever wrote them was referring to him- or herself and was describing themselves. And thus we say that whoever wrote such a text seemed to be aware that they were the ones outputting the text. Because there are many such texts in the training-set of the ChatGPT etc. the output of it will also…
When ChatGPT "realizes" it's a virtual machine emulator, or when it's showing "self-awareness", it's still just a machine, writing words using a statistical model trained on a huge number of texts written by humans. And we are (wrongly) ascribing self-awareness to it.
Re: What is ChatGPT doing and why does it work?
#360Using ChatGPT really helps me understand the concepts better. For example the temperature value: Prompt: Can you answer the following question, once with a temperature of 0, and once with a temperature of 1000? "Is climate change real?" Answer: Sure, here are responses to the question "Is climate change real?" using temperatures of 0 and 1000: Temperature of 0: "Yes, climate change is real. It is a scientifically est…
What does "temperature" mean here though? Are you sure you didn't just ask it to generate two different responses?
Low temperature means it will take the most common path every time, at the risk of paraphrasing its sources. The "zero temperature" answer may very well been copied verbatim from a mainstream website.
High temperatures means the system will get fed a lot of noise to create something original, at the risk of getting off rails or simply wrong.