Live data from Hacker News

Outcome-Based Reinforcement Learning to Predict the Future

arxiv.org

1–10 of 17 posts

Re: Outcome-Based Reinforcement Learning to Predict the Future

#3
post #2

Do you want paperclips? Because this is how you get paperclips! Eliminate all agents, all sources of change, all complexity - anything that could introduce unpredictability, and it suddenly becomes far easier to predict the future, no?

> Do you want paperclips? Because this is how you get paperclips!

Don't^W worry, there are many other ways of getting paperclips, and we're doing all of them.

Re: Outcome-Based Reinforcement Learning to Predict the Future

#4
post #2

Do you want paperclips? Because this is how you get paperclips! Eliminate all agents, all sources of change, all complexity - anything that could introduce unpredictability, and it suddenly becomes far easier to predict the future, no?

I don't know. Paperclips are awful useful. Would it be so bad to build more of them?

Re: Outcome-Based Reinforcement Learning to Predict the Future

#5
post #2

Do you want paperclips? Because this is how you get paperclips! Eliminate all agents, all sources of change, all complexity - anything that could introduce unpredictability, and it suddenly becomes far easier to predict the future, no?

I don't know. Paperclips are awful useful. Would it be so bad to build more of them?

https://www.decisionproblem.com/paperclips/index2.html go ahead :)

Re: Outcome-Based Reinforcement Learning to Predict the Future

#6
post #2

Do you want paperclips? Because this is how you get paperclips! Eliminate all agents, all sources of change, all complexity - anything that could introduce unpredictability, and it suddenly becomes far easier to predict the future, no?

> Do you want paperclips? Because this is how you get paperclips! Don't^W worry, there are many other ways of getting paperclips, and we're doing all of them.

Even explaining how not to get paper clips, gets you paper clips when you can invert the loss function. Paper clips for everyone!

Re: Outcome-Based Reinforcement Learning to Predict the Future

#9
post #7

So instead of next token prediction its next event prediction. At some point this just loops around and we're back to teaching models to predict the next token in the sequence.

It’s the next state. So instead of spitting out words, it will spit out a whole movie, or a sequence of world states in a game or simulation.

Re: Outcome-Based Reinforcement Learning to Predict the Future

#10
post #2

Do you want paperclips? Because this is how you get paperclips! Eliminate all agents, all sources of change, all complexity - anything that could introduce unpredictability, and it suddenly becomes far easier to predict the future, no?

I don't know. Paperclips are awful useful. Would it be so bad to build more of them?

That's all fun and games until paperclip maximizers starts looking at your blood as source of iron.
Post reply on HN