I was curious how good a transcription I could get from what may be the best multimoldal LLM currently, Gemini-1.5-Pro-Experiment-0801, so I had it transcribe five minutes of an interview between Ezra Klein and Nancy Pelosi from earlier today. The results are here: https://www.gally.net/temp/20240809geminitranscription/index... Aside from some minor punctuation and capitalization issues, Gemini’s transcription looks…
- auditory cues
- the sentence would be gramatically incorrect and make no sense without them
Just guessing out of the blue.
But I think it's likely that LLMs (and other speech recognition systems) need to exploit sentence context to recognize individual words and punctuation, and this is an example were it went well.
Human listening is similar in a way, we can recognize words even when spoken very mumbly or fast, if we have context.
So we always hear phrased rather than words.