Retentive Network: A Successor to Transformer for Large Language Models
21–22 of 22 posts
Re: Retentive Network: A Successor to Transformer for Large Language Models
#22Earlier quoted context omitted.
Not, but having a small section in the paper would be reasonable, that illustrates why the most pertinent models that might be relevant at first sight (like the ones I cited) are actually not applicable. The onus is on the authors to place their research in context and provide compelling arguments - not on the reader to guess why their model was compared against model A, but not model B. What do I mean by "pertinent"…
Critiques of omission are the easiest ones to levy and are informationally asymmetric in that the person leveling them will typically have more information on the topic being omitted than the person being critiqued. Hopfield networks & Metaformers are largely irrelevant in the domain of potential transformer replacements for LLM. Unless they have been shown to be relevant in any way, I don't see why the paper need ci…
Fair. Your argument then falls precisely into category (C) of the four mutually exclusive options I outlined above.
But you'd then need to argue why the 6 models you compared against is the comprehensive model sample to test against, that contains -- and not just some arbitrary set of recent models that happen to be dominated by the newly proposed model. (And maybe that is indeed the case; then it should be easy enough to update the arxiv draft by incorporating a section where you argue along those lines.)