Live data from Hacker News

The Revolution of Token-Level Rewards

levroai.com

1–6 of 6 posts

Re: The Revolution of Token-Level Rewards

#5
it was just a matter of time before the reward/error/etc. would become a vector instead of a scalar - whether it is a vector of tokens or a vector of some other qualities the output is evaluated for. So, instead of vector gradient - d scalar error by d every var - we'll have a matrix and thus would be matrix-transforming the NN instead of just adding the gradient.