No Train No Gain:Revisiting Efficient Training Algrthm for Transformer-BasedLM
1–2 of 2 posts
Re: No Train No Gain:Revisiting Efficient Training Algrthm for Transformer-BasedLM
#2Recent paper highlights the difficulty of creating a new optimizer as drop-in replacement. Sophia and Lion were recently proposed as superior alternatives to Adam, but appeared worse in an independent eval