Earlier quoted context omitted.
Wouldn't cost that much if the transcribing is done on device
This would be immediately obvious in a cursory analysis of performance. On-device transcription is not only computationally infeasible, it would also require model capabilities far beyond what is currently SOTA. Google had (and has afaik) significant challenges implementing multiple wake-word detection for precisely this reason. Transcribing a couple of words accurately on-device without a major performance penalty (…
Of course running it 24/7 in the background would ruin my battery, you would have to be smarter than that.