OpenAI Audio Models
11–20 of 317 posts
Re: OpenAI Audio Models
#12Does anyone have any experience with the realtime latency of these Openai TTS models? ElevenLabs has been so slow (much slower than the latency they advertise), which makes it almost impossible to use in realtime scenarios unless you can cache and replay the outputs. Cartesia looks to have cracked the time to first token, but i've found their voices to be a bit less consistent than Eleven Labs'.
Re: OpenAI Audio Models
#13Interesting, I inserted a bunch of "fuck"s in the text and the "NYC Cabbie" voice read it all just fine. When I switched to other voices ("Connoisseur", "Cheerleader", "Santa"), it responded "I'm sorry I can't assist with that request". I switched back to "NYC Cabbie" and it again read it just fine. I then reloaded the session completely, refreshed the voice selections until "NYC Cabbie" came up again, and it still r…
Re: OpenAI Audio Models
#14It would be much more convenient to use if changing the voice model worked on the fly, without having to stop and start the audio.
Re: OpenAI Audio Models
#15Re: OpenAI Audio Models
#16How do you call this desing/ui astethic? I like it
Re: OpenAI Audio Models
#17How do you call this desing/ui astethic? I like it
Re: OpenAI Audio Models
#18Re: OpenAI Audio Models
#19How do you call this desing/ui astethic? I like it
Re: OpenAI Audio Models
#20Nova+Serene sounds very metallic at the beginning about 50% of the time for me.
we put little stars in the bottom right corner for the newer voices, which should sound better