Live data from Hacker News

Why traditional reinforcement learning will probably not yield AGI [pdf]

philpapers.org

121–123 of 123 posts

Re: Why traditional reinforcement learning will probably not yield AGI [pdf]

#121

Earlier quoted context omitted.

"Necessarily" has general usage as well, you know... why would you read it otherwise, especially given the reasonable observation you make about this site? And my original point is not actually wrong, either: whether reinforcement learning will proceed at the pace of evolution is a topic of speculation - it is possible that it will, and possible that it will not. Insofar is I have an issue with your comment, it is th…

>> Insofar is I have an issue with your comment, it is that it is not going anywhere, as I explained in my previous post. I see this god-moding of my comment as a pretend-polite way to tell me I'm takling nonsense, that seems to be designed to avoid criticism for being rude to one's interlocutor on a site that has strong norms against that sort of thing, but without really trying to understand why those norms exist,…

"Necessarily", when read according to your own expectations for this forum, made an important difference to my original post (without it, I would have been insisting that the issue is settled already), so it was reasonable for me to point out its removal. The nitpicking over it began with your response to me doing so, and you have kept it going by taking the worst possible reading of what I write. This is, indeed, how things sometimes go.

Meanwhile, in a branching thread, I had a short discussion with the author of the post I originally replied to, in which I agreed with the points he made there. Both of us, I think, clarified our positions and reached common ground. That is how it is supposed to go.

I did not set out to pick a fight with you, and if I had anticipated how you would take my words, I would have phrased things more clearly.

Re: Why traditional reinforcement learning will probably not yield AGI [pdf]

#122
post #117

Earlier quoted context omitted.

It's not about whether the agent will or will not find an optimal solution (no agent can possibly find good solutions in every environment). It's about whether the agent could even understand the true environment based on your real-value-reward description of it. Suppose the true environment has various tasks the agent can do which give big-O-complexity-valued rewards, one task giving a reward of O(2^n), and others g…

Your paper does not prove anything in regards to what agent "mis-undestands". "Understanding" is not a mathematical concept in this case. It would still progressively find more and more optimal solutions to this task, so I see no problem there.

Suppose I claimed it's impossible to understand Shakespeare with all the consonants are deleted, leaving only vowels. Most people would accept that claim without pedantically asking me to define "understand". Same story here. I'm arguing that certain environments, when shoe-horned into the traditional RL model, are like Shakespeare with all the consonants removed.

The agent might (or might not) find progressively more and more optimal solutions to the interpreted environment, but not to the true environment (unless by dumb luck), because it does not see what the true environment actually is. For example, in the concrete environment described above, you could suppose completing Task C takes more steps than Task B. Then the agent is likely to conclude, on the basis of the misleading arctan rewards, that Task C is not worth the extra steps. Depending on the extra steps, this conclusion could be badly false.

Re: Why traditional reinforcement learning will probably not yield AGI [pdf]

#123
post #54
post #46

Earlier quoted context omitted.

do you know about the dopamine reward error hypothesis? https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6721851/ is it so wrong? what does cognitive psychology have to say about how these neurons work? this is a lot more recent than the 40s and behaviorism.

dopamine rewards operate on a different time scale vs. that required by these error correction models. I don't remember the exact paper, will need to look it up, but it was orders of magnitude difference in response times. Edit: for authoritative reference on biologically-plausible learning see anything by Edmund Rolls [1]. He explicitly stated in his recent book [2] that something like back-propagation, or similar e…

thanks for the link, it's been a lovely rabbit hole :)
Post reply on HN