Live data from Hacker News

The inefficiency of RL, and implications for RLVR progress

dwarkesh.com

11–20 of 50 posts

Re: The inefficiency of RL, and implications for RLVR progress

#12
post #5

Bit of a nitpick, but I think his terminology is wrong. Like RL, pretraining is also a form of *un*supervised learning

Usual terminology for the three main learning paradigms: - Supervised learning (e.g. matching labels to pictures) - unsupervised learning / self-supervised learning (pretraining) - reinforcement learning Now the confusing thing is that Dwarkesh Patel instead calls pretraining "supervised learning" and you call reinforcement learning a form of unsupervised learning.

SL and SSL are very similar "algorithmically": both use gradient descent on a loss function of predicting labels, human-provided (SL) or auto-generated (SSL). Since LLMs are pretrained on human texts, you might say that the labels (i.e., next token to predict) were in fact human provided. So, I see how pretraining LLMs blurs the line between SL and SSL.

In modern RL, we also train deep nets on some (often non trivial) loss function. And RL is generating its training data. Hence, it blurs the line with SSL. I'd say, however, it's more complex and more computationally expensive. You need many / long rollouts to find a signal to learn from. All of this process is automated. So, from this perspective, it blurs the line with UL too :-) Though it dependence on the reward is what makes the difference.

Overall, going from more structured to less structured, I'd order the learning approaches: SL, SSL (pretraining), RL, UL.

Re: The inefficiency of RL, and implications for RLVR progress

#13
post #6

In the limit, the "happy" case (positive reward), policy gradients boil down to performing more or less the same update as the usual supervised strategy for each generated token (or some subset of those if we use sampling). In the unhappy case, they penalise the model for selecting particular tokens in particular circumstances -- this is not something you can normally do with supervised learning, but it is unclear to…

The trick is to provide dense rewards, i.e. not only once full goal is reached, but a little bit for every random flailing of the agent in the approximately correct direction.

How do you know the correct direction? Isn’t the point of learning that the right path is unknown to start with?

Re: The inefficiency of RL, and implications for RLVR progress

#14
post #5

Earlier quoted context omitted.

Usual terminology for the three main learning paradigms: - Supervised learning (e.g. matching labels to pictures) - unsupervised learning / self-supervised learning (pretraining) - reinforcement learning Now the confusing thing is that Dwarkesh Patel instead calls pretraining "supervised learning" and you call reinforcement learning a form of unsupervised learning.

You could think of supervised learning as learning against a known ground truth, which pretraining certainly is.

a large number of breakthroughs in AI are based on turning unsupervised learning into supervised learning (alphazero style MCTS as policy improvers are also like this). So the confusion is kind of intrinsic.

Re: The inefficiency of RL, and implications for RLVR progress

#15
post #10

The premise of this post and the one cited near the start ( https://www.tobyord.com/writing/inefficiency-of-reinforcemen... ) is that RL involves just 1 bit of learning for a rollout, rewarding success/failure. However, the way I'm seeing this is that a RL rollout may involve, say, 100 small decisions out of a pool of 1,000 possible decisions. Each training step, will slightly upregulate/downregulate a given training…

It is the same type of learning, fundamentally: increasing/decreasing token probabilities based on the left context. RL simply provides more training data from online sampling.

Re: The inefficiency of RL, and implications for RLVR progress

#18
post #17

Earlier quoted context omitted.

[flagged]

[flagged]

There needs to be a new law, applicable to posts on the Internet of any kind.

Because that law doesn't hold, when malice has a massive profit motive, and almost zero downside.

Spammers, popups, spam, clickbait, all of it and more, not stupid, but planned.

Re: The inefficiency of RL, and implications for RLVR progress

#19

Earlier quoted context omitted.

Thank god. Was driving me mad.

[flagged]

That is a bizarre take. Dwarkesh Patel is publishing in a very specific domain, where RL is a very common and unambigous acronym. I'd bet it was immediately clear to 99% of his normal audience, and to him it's such a high frequency term that people finding it ambiguous would not even have crossed his mind.

(Like, would you expect people to expand LLM or AGI in a title?)

Re: The inefficiency of RL, and implications for RLVR progress

#20
post #13
post #6

Earlier quoted context omitted.

The trick is to provide dense rewards, i.e. not only once full goal is reached, but a little bit for every random flailing of the agent in the approximately correct direction.

How do you know the correct direction? Isn’t the point of learning that the right path is unknown to start with?

The correct solutions and the viable paths probably are known to the trainers, just not to the trainee. Training only on problems where the solution is unknown but verifiable sounds like the ultimate hard mode, and pretty hard to justify unless you have a model that's already saturated the space of problems with known solutions.

(Actually, "pretty hard to justify" might be understating it. How can we confidently extract any signal from a failure to solve a problem if we don't even know if the problem is solvable?)

Post reply on HN