Live data from Hacker News

OpenAI Audio Models

openai.fm

31–40 of 317 posts

Re: OpenAI Audio Models

#31

Quite disappointing their speech to text models are not open source. Whisper was really good and it was great it was open to play around with. I guess this continues OpenAI's approach of not really being open!

Indeed. Right now I think our open choices are Piper, Kokoro and Orpheus.

He was talking about STT models, not TTS. Whisper is open source and a good solution in many cases (in particular finetuned ones).

Re: OpenAI Audio Models

#32

Hey, I'm Jeff and I was PM for these models at OpenAI. Today we launched three new state-of-the-art audio models. Two speech-to-text models—outperforming Whisper. A new TTS model—you can instruct it how to speak (try it on openai.fm!). And our Agents SDK now supports audio, making it easy to turn text agents into voice agents. We think you'll really like these models. Let me know if you have any questions here!

Are the new models released with weights under an open license like whisper? If not, is it planned for the future?

Re: OpenAI Audio Models

#33
post #10

Interesting, I inserted a bunch of "fuck"s in the text and the "NYC Cabbie" voice read it all just fine. When I switched to other voices ("Connoisseur", "Cheerleader", "Santa"), it responded "I'm sorry I can't assist with that request". I switched back to "NYC Cabbie" and it again read it just fine. I then reloaded the session completely, refreshed the voice selections until "NYC Cabbie" came up again, and it still r…

Glad I'm not the only one whose inner 12 year old curiosity is immediately triggered by free input TTS. Swear words and just raking my hands across the keyboard to insert gibberish in every possible accent.

It won't generate slurs, though.

Re: OpenAI Audio Models

#35
post #7

Recommended input for anyone trying this out: Voice: Onyx Vibe: Heavy german accent, doing an Arnold Schwarzenegger impression, way over the top for comedic effect. Deep booming voice, uses pauses for dramatic effect.

interesting to see what accents do and do not work

Re: OpenAI Audio Models

#37

Hey, I'm Jeff and I was PM for these models at OpenAI. Today we launched three new state-of-the-art audio models. Two speech-to-text models—outperforming Whisper. A new TTS model—you can instruct it how to speak (try it on openai.fm!). And our Agents SDK now supports audio, making it easy to turn text agents into voice agents. We think you'll really like these models. Let me know if you have any questions here!

Are these models downloadable, like whisper?

What’s the minimum hardware for running them?

Would they run on a raspberry pi?

Or a smartphone?

Re: OpenAI Audio Models

#39

Hey, I'm Jeff and I was PM for these models at OpenAI. Today we launched three new state-of-the-art audio models. Two speech-to-text models—outperforming Whisper. A new TTS model—you can instruct it how to speak (try it on openai.fm!). And our Agents SDK now supports audio, making it easy to turn text agents into voice agents. We think you'll really like these models. Let me know if you have any questions here!

Is it against the TOS to use it for sexually explicit content?

I don't have your answer, but as far as innuendo goes, it's definitely capable!
Post reply on HN