The IPA itself isn't really sufficient to encode the sound of a word: you also need to know what language is being spoken; there is a little book you can get of vowel charts for all of the major languages in the world that tries to document the exact vocal position of each symbol. The issue is that there's a continuum of sound that can be generated by the human vocal system, and while humans don't want to differentia…
Thanks for taking the time to critique this - it's only something I put together in a few hours for fun. I'm sure somebody with more skill than me would be able to make something like this for multiple languages / accents.
Is this using parametric synthesis? (It doesn't sound concatenative.) Do you have a background in speech, signal processing, or audio, and is this just a passing interest, or something you want to continue to explore?
I've been teaching myself speech algorithms and methods off and on for the past six months ago. Recently I developed a concatenative Donald Trump text to speech engine (I've posted about it in the past), but the samples aren't great and it doesn't use proper unit selection. I'm trying to apply ML to generate a massive set of smooth n-phones that concatenate well together.
I'd definitely like to exchange contact info if you're into speech synthesis long term. My info is in my profile.
In any case, really cool project! :)