> MobileLLM-125M/350M attains a remarkable 2.7%/4.3% accuracy boost over preceding 125M/350M SoTA models on zero-shot commonsense reasoning tasks Small models, slightly improved, probably still not good enough for the same use as online models. Nothing wrong with incremental progress, however. 1.5B parameter model does seem to be a pretty decent step up, even beating larger models by a wide margin. I'm not sure why t…
An even smaller language model should still be useful as part of a speech-to-text system. These should benefit from using the language model to narrow down what word is spoken in the face of ambiguity or noise.