I just tested the "gpt-4o-mini-tts" model on several texts in Japanese, a particularly challenging language for TTS because many character combinations are read differently depending on the context. The produced speech was quite good, with natural intonation and pronunciation. There were, however, occasional glitches, such the word 現在
genzai “now, present” read with a pause between the syllables (
gen ...
zai) and the conjunction 而も read
nadamo instead of the correct
shikamo. There were also several places where the model skipped a word or two.
However, unlike some other TTS models offering Japanese support that have been discussed here recently [1], I think this new offering from OpenAI is good enough for language users. I certainly could have put it to good use when I was studying Japanese many years ago. But it’s not quite ready for public-facing applications such as commercial audiobooks.
That said, I really like the ability to instruct the model on how to read the text. In that regard, my tests in both English and Japanese went well.
[1] https://news.ycombinator.com/item?id=42968893