The Revolution of Token-Level Rewards
levroai.com
The Revolution of Token-Level Rewards
1–6 of 6 posts
Re: The Revolution of Token-Level Rewards
#2Publish a paper?
Re: The Revolution of Token-Level Rewards
#3Publish a paper?
We are working on one! Happy to send you a draft if you're interested.
Re: The Revolution of Token-Level Rewards
#4No implementation details, no samples from an actual reward model in action, no github repo. Looks like a sales page more than anything. Eww.
Re: The Revolution of Token-Level Rewards
#5it was just a matter of time before the reward/error/etc. would become a vector instead of a scalar - whether it is a vector of tokens or a vector of some other qualities the output is evaluated for. So, instead of vector gradient - d scalar error by d every var - we'll have a matrix and thus would be matrix-transforming the NN instead of just adding the gradient.