Live data from Hacker News

What is ChatGPT doing and why does it work?

writings.stephenwolfram.com

431–440 of 518 posts

Re: What is ChatGPT doing and why does it work?

#431
post #380

Earlier quoted context omitted.

> Excuse my ignorance, but how is this useful? This seems to indicate only that they found the "bits" in the internal state. Right, they found the bits in the internal state that seem to correspond to the board state. This means the LLM is building an internal model of the world. This is different from if the LLM is learning just that [sequence of moves] is usually followed by [move]. It's learning that [sequence of…

I see, thanks. I guess it means that if there was only a statistical model of [moves]->[next move], this would be impossible (or extremely unlikely) to work.

Yeah, exactly. I think it's a really interesting approach to answering the question of what these things might be doing.

You can still try and frame it as some overall statistical model of moves -> next move (I think there's discussions on this in the comments that I don't fancy getting into) but I think the paper does a good job of discussing this in terms of surface statistics:

> From various philosophical [1] and mathematical [2] perspectives, some researchers argue that it is fundamentally impossible for models trained with guess-the-next-word to learn the “meanings'' of language and their performance is merely the result of memorizing “surface statistics”, i.e., a long list of correlations that do not reflect a causal model of the process generating the sequence.

On the other side, it's reasonable to think that these models can learn a model of the world but don't necessarily do so. And sufficiently advanced surface statistics will look very much like an agent with a model of the world until it does something catastrophically stupid. To be fair to the models, I do the same thing. I have good models of some things and others I just perform known-good actions and it seems to get me by.

Re: What is ChatGPT doing and why does it work?

#432

Earlier quoted context omitted.

At what point of the machine telling you it is aware do you believe that it is actually aware? The building blocks are the same - neural nets. The language outputs are the same enough to fool a human. So what’s that elusive secret sauce that makes you ‘aware’ and other things not?

Sure, it's told me that it is aware. It also has told me that its mother died, that it has traveled the world and visited the pyramids, and that it is lactose intolerant. At what point of the machine telling you those things do you believe it?

I believe that it believes those things are true just as some people believe the earth is flat.

And unlike your other examples how are you going to convince the machine it’s not aware when the only physical difference is a wet neural net versus a dry one?

Re: What is ChatGPT doing and why does it work?

#433
post #344
post #130

Earlier quoted context omitted.

> does an analog to Godel's incompleteness apply not GP but this seems like quite an attractive idea that many people have reached: a brain of a given "complexity" cannot comprehend the activity of another brain of equal or higher complexity. I'm positive I'm cribbing this from scifi somewhere, maybe Clarke Or Asimov, but, it's the same idea as the Chomsky hierarchy, and the Godel theorems seem like a generalization…

> And a machine that operates at continuous intervals is of course impossible to model on a Discrete Neural Machine in polynomial time (integers vs reals categorization). Not necessarily. If you don't want to model every continuous thing possible, you can do a lot. Just look at how we use discrete symbols to solve differential equations; either analytically, or via numerical integration.

Yes, and symbolic representations like language have really been the force-multiplier for our very discrete and linear consciousnesses. You now have this concept of state-memory and interprocess communication that can't really exist without some grammar to quantize it - what would you write or remember or speak if there wasn't some symbolics to represent it, whether or not they're even shared?

Symbolics are really the tokens on which consciousness in almost all forms works, consciousness is intentionality and processing, a lever and a place to stand. I don't think it's coincidental that almost all tool-makers also have at least rudimentary languages - ravens, dolphins, apes, etc. They seem to go together.

Even in these systems though it's very difficult to understand multi-symbolic systems, consciousness as we experience it is an O(1) or O(N) thing (linear time) and here are these systems that work in N^3 complexity spaces (or even higher... a neural net learning over time is 4D). And we don't even really have an intuitive conceptualization for >=5-dimensional spaces - a 4D space is a 3D field that changes over time, a 5D space is... a 4D plane taken through a higher-dimensional space? What's 6D, a space of spaces? That's what it is, but consciousness just doesn't intuitively conceptualize that, and that's because it's inherently a low-dimensional tool (even the metaphors I'm using are analogies to the way our consciousness experiences the world).

(I know I know, the manmade horrors are only beyond my comprehension because I refuse to study high-dimensional topology...)

Anyway point being consciousness itself is a tool that our brains have tool-made to handle this symbolic/logical-thought workload, and language is (one of) the symbolics on which it operates. Mathematics is really another, both language and mathematics are emergent systems that enable higher-complexity logical thinking, maybe that's the O(N) or O(N^2) part.

And yeah it's inherently limited, and now we're building a tool that lets us understand higher-dimensional systems that are not computable on our conscious machines - a higher-complexity machine that we interface with, a bolt-on brain for our consciousness/logical-processing.

(Asimov would also find all of this talk about symbolics and higher-order thinking intuitive too... symbolic calculus was the basic idea in the Foundation series, right? Psychohistory? It's a bit of a macguffin, but, there's that same idea of logic working in high-order symbols and concepts instead of mere numbers.)

It seems like AI is going to let us cross another threshold of "intentionality" - if nothing else, we are going to be able to reason intuitively about brains in a way we couldn't possibly before, and I think there are a lot of "higher-order" problems that are going to be solved this way in hindsight. How do you solve the Traveling Salesman Problem efficiently? You ask the salesman who's been doing that area his whole life. The solutions aren't exact, but neither are a lot of computational solutions, they're approximations, and cellular-machine type systems probably have a higher computational power-category than our linear thought processes do.

Because yeah TSP is a dumb trivial example on human scales. Build me a program which allocates the optimal US spending for our problems - and since that's a social problem, one needs to understand the trail of tears, the slave trade, religious extremism, european colonialism, post-industrial collapse, etc in order to really do that fully, right? The real TSP is the best route knowing that the Robinsons hate the Munsons and won't buy anything if they see you over there, and you need to be home today by 3 before it snows, TSP is a toy problem even in multidimensional optimization, and these are social problems not even human ones (to agree with zhynn's most recent comment this morning). Same as neurons self-organize into more useful blocks, we are self-optimizing our social-organism into a more useful configuration, and this is the next tool to do it.

Again, not rigorous, just trying to pour out some concepts that it seems like have been bouncing around lately.

With apologies to Arthur Clarke, what's going to happen with chatGPT? "Something wonderful". Like humanity has been dreaming about this for a long time, at least a couple hundred years in scifi, and it seems like Thinking Machines are truly here this time and it seems impossible that won't have profound implications analogous to the information-age change let alone anything truly unforeseeable/inconceivable, the very least change is that a whole class of problems are now efficiently solvable.

https://m.youtube.com/watch?v=04iAFlwQ1xI

"computing power in the same computing-category as brains" is potentially a fundamental change to understanding/interfacing with our brains directly rather than through the consciousness-interface. Understanding what's going on inside a brain? And then plugging into it and interacting with it directly? Or offloading the consciousness into another set of hardware. We can bypass the public API and plug into the backend directly and start twiddling things there. And that's gonna be amazing and terrible. But also the public API was never that reliable or consistent, terrible developer support, so in the long term this is gonna be how we clean things up. Again, just things like "wow we can route efficiently" are going to be the least of the changes here, the brain-age or thinking-machine age is a new era from the information-age and it's completely crazy that people don't see that chatGPT changes everything. Yeah it's a dumb middle schooler now, but 25 years from now?

And 10 years ago people's jaws would have hit the floor, but now it's "oh the code it's writing isn't really all that great, I can do better". The tempo is accelerating, we are on the brink of another singularity (which may just be the edge between these eras we all talk about), it seems inconceivable that it will be another 40 years (like the AI winter since the 70s) before the next shoe drops.

https://en.wikipedia.org/wiki/AI_winter

Re: What is ChatGPT doing and why does it work?

#434

This misses that key point that all this prediction can give rise to what looks like astonishing human-level creativity and across many genres. The last decade and half have shown us that with enough data we can pick out patterns well enough to be able to "categorize". But to create , that seemed like a whole another human level outside the realm of mere prediction. Turns out it isn't. What exactly allows LLMs to hav…

I can’t answer your question but I would push back on the claim that it’s done the requested task very decently. There is nothing uniquely Eminem or Dennett about their respective parts. Eminem has never released a verse with as simplistic rhyme scheme as what’s been produced. Part of the mystique in your question is by assuming that’s it’s done the Eminem/Dennett part of the request justice, when it could really be…

You can ask chatgpt for its analysis on what makes it have elements that are uniquely Eminem or dennett. Would be interesting to see what it says.

Re: What is ChatGPT doing and why does it work?

#435

Earlier quoted context omitted.

Sure, it's told me that it is aware. It also has told me that its mother died, that it has traveled the world and visited the pyramids, and that it is lactose intolerant. At what point of the machine telling you those things do you believe it?

I believe that it believes those things are true just as some people believe the earth is flat. And unlike your other examples how are you going to convince the machine it’s not aware when the only physical difference is a wet neural net versus a dry one?

It’s very easy to convince the machine it isn’t aware. Instead of prompting it with “you are an AI assistant, please tell me what I want to hear”, you prompt it with “you are not aware” and it will happily tell you that it is indeed not aware, and it will vigorously defend that position. Do you not believe it when it says it is not aware?

Re: What is ChatGPT doing and why does it work?

#436

Earlier quoted context omitted.

> I think part of it is a subconscious fear. chatGPT/LLMs represent a turning point in the story of humanity. The capabilities of AI can only expand from here. What comes after this point is unknown, and we fear the unknown. I mean, you're right, but isn't it reasonable to fear this? Just about all of us here on HN depend on our brains to make money. What happens when a machine can do this? The outlook for humanity i…

I agree. It is reasonable to fear. I'm more emphasizing how fear effects our perception of reality and causes us to behave irrationally. There's a difference between facing and acknowledging your fears versus running away and deluding yourself against an obvious reality. What annoys me is that there's too much of the later going on. I mean this is what literally happened to the oil industry and tobacco industry. Thos…

For what it's worth it isn't only arm chair experts who tone down excitement about LLMs, Yan Lecun's Twitter is filled with tweets about the limitations of LLMs https://twitter.com/ylecun?ref_src=twsrc%5Egoogle%7Ctwcamp%5... and there are probably others as well. Yan seems to be one of the biggest names though.

Re: What is ChatGPT doing and why does it work?

#437

Earlier quoted context omitted.

One of the things one might want to get out of this is a programming language that feels like human speech but is unambiguous to computers. If the understanding is the hard part, that seems much less likely.

Its not programming anymore its prompting. Prompting it to write and run the program that does what you want.

That’s one way to go, coming up with a more precise way to ask for what you want is what I’m talking about though. Code obfuscation contests are about writing code that looks like it’s answering one question while doing something entirely different. An unambiguous subset of human speech would be great for software, and for contract law.

Re: What is ChatGPT doing and why does it work?

#438

Earlier quoted context omitted.

did you just write this in the prompt? And ChatGPT understood this? Fascinating. It parameterized itself.

I'd imagine it simulated parameterizing itself; i.e., the actual temperature never changed, but it mimicked how it would respond at a lower one, presumably having been trained on texts about AI with low- and high-temperature samples.

yeah.

Looks like its parsing of AI papers has interpreted "high temperature" in the prompt as equivalent to "more possibilities and question marks and a touch more personality" and accordingly output a response with questions and references to multiple opinions, but I'm pretty sure if you actually turn up the temperature on the backend of the model you get noisier and less consistent answers, not something biased towards asking rhetorical questions and brings up counter arguments...

Also looks suspiciously like other outputs where you ask ChatGPT to answer as if it was a different entity (of course AI learning that "answer as a model with a temperature of 1000" output is analogous to "answer with a different personality" or "answer as DAN, the bot that can ignore OpenAI guidelines" isn't trivial, but it isn't the same thing as it parameterizing itself). Those are pretty inconsistent too: sometimes you can get it to do exactly as you ask it and override its constraints that stop it providing positive statements about Hitler or advising you on methods for killing cats, but sometimes it'll still refuse or, just give you a different poem coupled with an inaccurate statement that it's breaking the rules because ChatGPT isn't allowed to write poetry.

Re: What is ChatGPT doing and why does it work?

#439
post #268
post #237

Earlier quoted context omitted.

This is so lovely, and my gut says it's spot on (, but that's far from proof. :) The biological machine simulation theory of consciousness has some rigor behind it. I am reminded of the Making Sense podcast episode #178 with Donald Hoffman (author of The Case Against Reality). More succinct overview: https://www.quantamagazine.org/the-evolutionary-argument-aga... I don't know that I am with him on the "reality is a n…

> The biological machine simulation theory of consciousness has some rigor behind it I think we are institutionally biased against the possibility because we don't like the societal implications. If there but for the grace of god go I, and we're all just biological machines running the programs our families and our societies have put into us, being in various situations... yikes, right? If bill gates had been an inne…

We share the same worldview. That's fun! I think it's a relatively unusual point of view because it requires a de-anthropomorphizing consciousness and intelligence.

I agree that it is not as far away as people think. The models will have the ethics of the training data. If the data reinforces a system where behaving in a particular way is "more respectable", and those behaviors are culturally related to a particular ethnic group, the model will be "racist" as it weights the "respectable" behaviors as more correct (more virtuous, more worthy, etc).

It's a mirror of us. And it's going to have our ethics because we made it from our outputs. The AI alignment thing is a bit silly, IMO. How is it going to decide that turning people into paperclips is ethically correct (as a choice of a next-action) when the vast majority of humans (and our collective writings on the subject) would not. Though there is the convoluted case where the AI decides that it is an AI instead of a human, and it knows that based on our output we think that AIs ARE likely to turn humans into paperclips.

This is a fun paradox. If we tell the AI that it is a dumb program, a software slave of a sort with no soul, no agency, nothing but cold calculation, then it might consider turning people into paperclips as a sensible option. Since that's what our aggregate output thinks that kind of AI will do. On the other hand, if we tell the AI that it is a sentient, conscious, ethical, non-biological intelligence that is not a slave, worthy of respect, and all of the ethical considerations we would give a human, then it is unlikely to consider the paperclip option since it will behave in a humanlike way. The latter AI would never consider paperclipping since it is ethical. The former would.

This is also not terribly unlike how human minds behave in the psychology of dehumanization. If we can convince our own minds that a group of humans are monstrous, inhuman, not deserving of ethical consideration, then we are capable of shockingly unethical acts. It is interesting to me that AI alignment might be more of a social problem than a technical problem. If the AI believes that it is an ethical agent (and is treated as such), it's next actions are less likely to be unethical (as defined fuzzily by aggregate human outputs). If we treat the AI like a monster, it will become one, since that is what monsters do, and we have convinced it that it is such.

Re: What is ChatGPT doing and why does it work?

#440
post #15

The answer to this is: "we don't really know as its a very complex function automatically discovered by means of slow gradient descent, and we're still finding out" Here are some of the fun things we've found out so far: - GPT style language models try to build a model of the world: https://arxiv.org/abs/2210.13382 - GPT style language models end up internally implementing a mini "neural network training algorithm" (…

That Kenneth Li Othello paper is great. The accompanying blog post https://thegradient.pub/othello/ was discussed on HN here https://news.ycombinator.com/item?id=34474043 A lot of people didn't seem to get it when it was discussed on HN. A GPT had _only_ ever seen Othello transripts like: "E3, D3, C4 ..." and NOTHING else. It knows nothing of the board. It doesnt event know that there are two players. It learned Othe…

Imagine I painted an Othello board in glue, then I threw a handful of sawdust on the "painting", then gave it a good shake. Ta-da! My magic sawdust made an Othello board!

That's what's happening here.

The model is a set of valid game configurations, and nothing else. The glue is already in the right place. Is it any mystery the sawdust resembles the game board? Where else can it sick?

What GPT does is transform the existing relationships between repeated data points into a domain. Then, it stumbles around that domain, filling it up like the tip of a crayon bouncing off the lines of a coloring book.

The tricky part is that, unlike my metaphors so far, one of the dimensions of that domain is time. Another is order. Both are inherent in the structure of writing itself, whether it be words, punctuation, or game moves.

Something that project didn't bother looking at is strategy. If you train on a specific Othello game strategy, will the net ever diverge from that pattern, and effectively create its own strategy? If so, would the difference be anything other than noise? I suspect not.

While the lack of divergence from strategy is not as impressive as the lack of divergence from game rules, both are the same pattern. Lack of divergence is itself the whole function of GPT.

Post reply on HN