I have a project idea that might be useful for someone to learn about audio processing (and maybe neural networks?). I like to listen to audio and watch video of podcasts (and lectures and other human speech) at faster speeds. Sometime, especially if I'm trying to "skim" to see if the media is worth listening to carefully, I'd like to listen at 3x or faster. Very often, the limiting factor is the intelligibility of the actual words rather than mentally parsing them.
Some software already removes complete silences, but this is a 10% effect and I think this could be taken much further. I would love audio software that could manipulate high-speed human speech to improve intelligibility by preferentially compressing parts with low information content (like vowels and "ughs") and uncompressing, or even "repairing", info-dense parts like sequential consonant sounds.
I've looked around and haven't been able to find anything like this. Could make a nice stand-alone app, or a library to sell to a podcast player.
http://softwarerecs.stackexchange.com/questions/27175/video-...