Live data from Hacker News

Why traditional reinforcement learning will probably not yield AGI [pdf]

philpapers.org

71–80 of 123 posts

Re: Why traditional reinforcement learning will probably not yield AGI [pdf]

#71
post #17
post #13

Earlier quoted context omitted.

Broken in the sense that it's not flexible enough to apply to all conceivable environments a genuine AGI could navigate, without misleading that AGI. But I should stress that real number objective functions are probably fine for many specific interesting environments, I'm not trying to say that real number objective functions are useless. Just that they aren't flexible enough to cover all environments :)

Fair enough. It would be interesting/instructive to construct (relatively simple) examples where we can see that they’re broken :-)

Easy to come up with examples using exotic money-related constructions. Suppose there's something called a "superdollar". If you have a superdollar, you can use it to create an arbitrary number of dollars for yourself, any time you want, which you can trade for goods and services. If you want, you can also trade the superdollar itself. Now picture an environment with two buttons, one of which always rewards you one dollar, and the other of which always rewards you one superdollar. Shoe-horning this environment into traditional RL, you'd have to assign the superdollar button some finite reward, say a million. But then you would mislead the traditional-RL-agent into thinking a million dollars was as good as one superdollar, which clearly is not true.

Re: Why traditional reinforcement learning will probably not yield AGI [pdf]

#72
post #7

Earlier quoted context omitted.

The argument isn't so much about the type of agent (which I think is what Neural Turing Machines etc. are about), it's about the type of environment. In traditional Reinforcement Learning, environments give real-number-valued rewards (or even rational-number-valued rewards which is even more constrained). Presumably this was a decision that was made with hardly a second thought because real numbers are most familiar…

But isn't the sophisticated structure what leads to the real number. If you change the rewards to something more complex surely you still have to pick between actions and at some point you'll have to evaluate which one is "better" and I can't see why you couldn't use real numbers to represent utility. I mean, humans are general intelligences, and you can translate pretty much any human reward into money, which is a r…

Question: when you say "I can't see why you couldn't use real numbers to represent utility", does your reasoning for that have anything to do with Dedekind cuts, Cauchy sequences, or complete ordered fields? Because that's what the real numbers _are_. If your reasoning has nothing to do with these sort of things, then it can't possibly be sound because in order to argue that X has such-and-such property, you need to know what X actually _is_.

To repeat an example I posted for someone else: Suppose there's something called a "superdollar". If you have a superdollar, you can use it to create an arbitrary number of dollars for yourself, any time you want, which you can trade for goods and services. If you want, you can also trade the superdollar itself. Now picture an environment with two buttons, one of which always rewards you one dollar, and the other of which always rewards you one superdollar. Shoe-horning this environment into traditional RL, you'd have to assign the superdollar button some finite reward, say a million. But then you would mislead the traditional-RL-agent into thinking a million dollars was as good as one superdollar, which clearly is not true.

Re: Why traditional reinforcement learning will probably not yield AGI [pdf]

#73
post #12

Earlier quoted context omitted.

>Humans don't receive infinite rewards Maybe you're right, but I don't think it's as obvious as you imply. Consider this question. Can there exist hypothetical rewards x1,x2,x3,..., each one of which is significantly better than the previous, and at the same time another hypothetical reward y such that y is significantly better than each x_i? If such rewards can exist, and if "significantly better" implies "at least…

Humans don't receive infinite rewards because any "reward system" in our brain is implemented by receiving a finite amount of electrochemical pleasure signals over a finite amount of time. Human-equivalent minds can obviously be implemented on top a framework that does not have inherently infinite or infinitely divisible values, because human minds are implemented on a top of a substrate that uses a finite amount of…

>Humans don't receive infinite rewards because any "reward system" in our brain is implemented by receiving a finite amount of electrochemical pleasure signals over a finite amount of time.

This is like saying computers can't represent infinity because they have only finitely many bytes.

Suppose the treasury rewarded you a "superdollar", which is a special object that allows you to create any number of dollars that you want, on demand, as many times as you want. How many dollars would you say this superdollar is worth? Obviously, no finite number of dollars would be worth that one superdollar. The human mind can certainly understand the relative value of a superdollar vs. any number of dollars. That the human mind is implemented through finite electrochemical processes is irrelevant.

Re: Why traditional reinforcement learning will probably not yield AGI [pdf]

#74
I like reductionist maths counter-proofs based on "x cannot contain y", but I try to be skeptical they apply, because it is plain real mathematicians can reason about infinite things, from finite symbols, as chains of symbols. the "cannot contain" set includes things I can state, but not enumerate.

I'm skeptical about AGI anyway. The proof is unnecessary, I tend to an even more reductionist model: Our lack of understanding where intelligence is, in the brain, goes to our failure to model it. "morally" making claims "its alive" based on a lack of understanding, is a bit like chemistry by alchemy. If you don't know why, you didn't make it any more than making a baby would have.

Feynman said it better: lots of physics proofs are built on partial models of sub-atomics, which if you ask questions about become unknowables too. Its turtles-all-the-way-down stuff.

traditional reinforcement learning will probably not yield AGI, any more than any current method, based on not understanding GI, and until we understand GI, I do not believe any connectionist, or learning model will derive it.

Re: Why traditional reinforcement learning will probably not yield AGI [pdf]

#75
post #19

I applaud the effort but the problem with RL as a model of learning is in the definition of RL itself. The idea of using "rewards" as a primary learning mechanism and a path to actual cognition is just wrong, full stop. It's a wrong level of abstraction and is too wasteful in energy spent. Looking at it from CogSci perspective it is essentially an offshoot of behaviorism, using a coarse and extremely inefficient mode…

[deleted]

Re: Why traditional reinforcement learning will probably not yield AGI [pdf]

#76

Earlier quoted context omitted.

Except no organism is born a blank slate. Parent is correct in that our prior was massively expensive to construct

So we can expect our ANN’s to yield AGI in a few million or billion years? That doesn’t sound like a good place to put our current efforts then.

That does not necessarily follow, as I imagine you well know.

Re: Why traditional reinforcement learning will probably not yield AGI [pdf]

#77
post #66

Earlier quoted context omitted.

I think the point huh is being mad is that individual people (or models) dot learn that way. It’s not like models training models, all the way down.

Individual people are not trained from scratch. ML models often have to be (modulo fine-tuning) since the field is still young.

That's already changing. That we have only relatively recently moved beyond always starting from scratch might indicate that we are still in the Cambrian of AI, however...

Re: Why traditional reinforcement learning will probably not yield AGI [pdf]

#78
post #52
post #2

Author here. One particularly topical observation (topical because HN has recently featured discussions about ways to merge statistical and symbolic approaches to AI), from Section 4.2.3: certain cutting-edge number systems such as Conway's "surreal numbers" are so sophisticated that they require lots of symbolic logic just to do basic operations with. Of course this makes these number systems hard to work with, but…

do your claims apply in complex-valued RL context?

I'm surprised to see that that's actually a thing people write about. I'm afraid I can't answer your question right now, as I have no idea why anyone would do that (as opposed to, say, vector-valued rewards---i.e., why is the ability to complex-multiply two complex rewards relevant, as opposed to merely comparing separate components of rewards from R^2). I'll have to read up on the subject :)

Re: Why traditional reinforcement learning will probably not yield AGI [pdf]

#79
post #19

I applaud the effort but the problem with RL as a model of learning is in the definition of RL itself. The idea of using "rewards" as a primary learning mechanism and a path to actual cognition is just wrong, full stop. It's a wrong level of abstraction and is too wasteful in energy spent. Looking at it from CogSci perspective it is essentially an offshoot of behaviorism, using a coarse and extremely inefficient mode…

>state-of-the-art cognitive psychology, and may be looking at research in "distributional semantics", "concept spaces", "sparse representations", "small-world networks" and "learning and memory" neuroscience.

Look, uh, I've read Gardenfors too, but are those really the state of the art? I don't remember there being anything about them at CogSci this past summer. Maybe I wasn't paying close-enough attention?

Re: Why traditional reinforcement learning will probably not yield AGI [pdf]

#80
post #19

I applaud the effort but the problem with RL as a model of learning is in the definition of RL itself. The idea of using "rewards" as a primary learning mechanism and a path to actual cognition is just wrong, full stop. It's a wrong level of abstraction and is too wasteful in energy spent. Looking at it from CogSci perspective it is essentially an offshoot of behaviorism, using a coarse and extremely inefficient mode…

Reinforcement learning is Turing complete [1], so if AI is possible at all, then it can be realised through RL. cognitive psychology You are overselling the insights of this discipline. Has cognitive psychology solved its replication problems? Where is the world-beating AI that is based on "concept spaces", "sparse representations", "small-world networks" and "learning and memory" neuroscience? [1] https://arxiv.org/…

>Reinforcement learning is Turing complete [1], so if AI is possible at all, then it can be realised through RL.

Brainfuck is also Turing complete, so logically if we just do, for instance, Markov chain Monte Carlo for Bayesian program learning in Brainfuck, we can realize AGI that way.

"Everything is possible, but nothing is easy." The Turing tarpit.

Post reply on HN