Live data from Hacker News

Outcome-Based Reinforcement Learning to Predict the Future

arxiv.org

11–17 of 17 posts

Re: Outcome-Based Reinforcement Learning to Predict the Future

#11
post #8
post #7

So instead of next token prediction its next event prediction. At some point this just loops around and we're back to teaching models to predict the next token in the sequence.

Tokens are an awfully convenient way to describe an event.

Tokens are just discretized state representations.

Re: Outcome-Based Reinforcement Learning to Predict the Future

#14

From the abstract > A simple trading rule turns this calibration edge into $127 of hypothetical profit versus $92 for o1 (p = 0.037). I'm lazy: is this hypothetical shooting fish in a barrel, or is it a real edge?

Note the 'hypothetical profit' part , I know of several groups looking for opportunities to skim off LLM traders, leveraging its limited sensitivity, expressiveness, and the loss of tail data.

Predictive AI is problematic no matter what tool you use. Great at demoware that doesn't deliver.

I am sure there are use cases, but it would be augmentation, not a reliable approach by itself.

Post reply on HN