So instead of next token prediction its next event prediction. At some point this just loops around and we're back to teaching models to predict the next token in the sequence.
Tokens are an awfully convenient way to describe an event.
Outcome-Based Reinforcement Learning to Predict the Future
11–17 of 17 posts
Re: Outcome-Based Reinforcement Learning to Predict the Future
#12Re: Outcome-Based Reinforcement Learning to Predict the Future
#13> A simple trading rule turns this calibration edge into $127 of hypothetical profit versus $92 for o1 (p = 0.037).
I'm lazy: is this hypothetical shooting fish in a barrel, or is it a real edge?
Re: Outcome-Based Reinforcement Learning to Predict the Future
#14From the abstract > A simple trading rule turns this calibration edge into $127 of hypothetical profit versus $92 for o1 (p = 0.037). I'm lazy: is this hypothetical shooting fish in a barrel, or is it a real edge?
Predictive AI is problematic no matter what tool you use. Great at demoware that doesn't deliver.
I am sure there are use cases, but it would be augmentation, not a reliable approach by itself.
Re: Outcome-Based Reinforcement Learning to Predict the Future
#15So instead of next token prediction its next event prediction. At some point this just loops around and we're back to teaching models to predict the next token in the sequence.
Re: Outcome-Based Reinforcement Learning to Predict the Future
#16Why would you use RL if you're not going to control the environment, but just predict it?
Re: Outcome-Based Reinforcement Learning to Predict the Future
#17bzzzzz "sorry this isn't your lucky day"