I used to work for a company that built Speech Recognition systems and I came up with a similar/related idea - the idea being to take a load of videos of barack obama (for example), and create an accurate 'voice print'. Once done, any videos or speech could be scanned and if Barack Obama's voice print was recognized/detected, the recognizer could be tuned to his voice print AND could apply a set of appropriate grammars/vocabluary (for example the 'politics' grammar, or 'american' grammar or 'economics' grammar) - then you could very accurately perform speech recognition and automatically create text translations. Then when you google for text, you could actually retrieve videos whose content exactly matches the search terms and jump directly to that part of the video.
Over time you could build up a database of voice prints and grammars for not just celebrities, policitians, but also criminals (for automatic identification).
I had this idea almost 4 years ago, submitted it to the company, but it wasn't taken seriously.
If anybody is interested in this, let me know!