Live data from Hacker News

OpenAI Audio Models

openai.fm

121–130 of 317 posts

Re: OpenAI Audio Models

#121

Earlier quoted context omitted.

Any plans to directly support diarization or voiceprinting?

We're thinking about diarization (adding time awareness to GPT models) but no firm plans to share just yet

The feature I want is speaker differentiation - I want to feed in an audio file and get back a transcript with "Speaker 1: ..., Speaker 2: ..." indications.

That plus timestamps would be incredible.

The Google Gemini 2.0 models are showing some promise with this, I can't speak to their reliability just yet though.

Re: OpenAI Audio Models

#122

Earlier quoted context omitted.

> Two speech-to-text models—outperforming Whisper On what metric? Also Whisper is no longer state of the art in accuracy, how does it compare to the others in this benchmark? https://artificialanalysis.ai/speech-to-text

We've been using the FLUERS eval and you can see comparisons to other models on the market in the post https://openai.com/index/introducing-our-next-generation-aud... Curious if there's a benchmark you trust most?

FLUERS and GP's Common Voice dataset focus on read speech. I've observed models that perform well on these datasets be completely useless on other distributions, like whispered speech or shouted speech or conversational speech between humans who aren't talking to a computer.

Re: OpenAI Audio Models

#123
post #37

Earlier quoted context omitted.

Are these models downloadable, like whisper? What’s the minimum hardware for running them? Would they run on a raspberry pi? Or a smartphone?

not open source at this time. unfortunately they're much to large to run on normal consumer hardware

[deleted]

Re: OpenAI Audio Models

#124

Hey, I'm Jeff and I was PM for these models at OpenAI. Today we launched three new state-of-the-art audio models. Two speech-to-text models—outperforming Whisper. A new TTS model—you can instruct it how to speak (try it on openai.fm!). And our Agents SDK now supports audio, making it easy to turn text agents into voice agents. We think you'll really like these models. Let me know if you have any questions here!

[deleted]

Re: OpenAI Audio Models

#125

Hey, I'm Jeff and I was PM for these models at OpenAI. Today we launched three new state-of-the-art audio models. Two speech-to-text models—outperforming Whisper. A new TTS model—you can instruct it how to speak (try it on openai.fm!). And our Agents SDK now supports audio, making it easy to turn text agents into voice agents. We think you'll really like these models. Let me know if you have any questions here!

Woohoo new voices! I’ve been using a mix of TTS models on a project I’ve been working on, and I consistently prefer the output of OpenAI to ElevenLabs (at least when things are working properly).

Which leads me to my main gripe with the OpenAI models — I find they break — produce empty / incorrect / noise outputs — on a few key use cases for my application (things like single-word inputs — especially compound words and capitalized words, words in parenthesis, etc.)

So I guess my question is might gpt-4o-mini-tts provide more “reliable” output than tts-1-hd?

Re: OpenAI Audio Models

#126
post #7

Recommended input for anyone trying this out: Voice: Onyx Vibe: Heavy german accent, doing an Arnold Schwarzenegger impression, way over the top for comedic effect. Deep booming voice, uses pauses for dramatic effect.

This is hilarious, extra points if you get it to say:

"Get to the chopper now and PUT THAT COOKIE DOWN NOWWWW"

Re: OpenAI Audio Models

#127
post #7

Recommended input for anyone trying this out: Voice: Onyx Vibe: Heavy german accent, doing an Arnold Schwarzenegger impression, way over the top for comedic effect. Deep booming voice, uses pauses for dramatic effect.

so good! https://www.openai.fm/#28540f27-5b51-445a-b1d6-1c89711a2c4f

Weird 1 in maybe 5 times when I refresh that page I get a very good Arnold. The rest are very bad.

Re: OpenAI Audio Models

#128
post #65

Earlier quoted context omitted.

ElevenLabs is incredibly over-priced and that's how they were able to achieve the MRR that led to their incredible fundraising. No matter what happens, they'll eventually be undercut and matched in terms of quality. It'll be a race to the bottom for them too. ElevenLabs is going to have a tough time. They've been way too expensive.

I hope they find a more unique product offering that takes hold. Everybody thinks of them as text-to-speech but I use ElevenLabs exclusively for speech-to-speech for vtubing as my AI character. They're kind of the only game in town for doing super high quality speech-to-speech (unless someone here has an alternative which I'd LOVE to know about). I've tried https://github.com/w-okada/voice-changer which is great beca…

Are you comfortable sharing the video & lip-sync stack you use? I don't know anything about the space but am curious to check out what's possible these days.

Re: OpenAI Audio Models

#129
post #29

Earlier quoted context omitted.

Indeed. Right now I think our open choices are Piper, Kokoro and Orpheus.

In my opinion GPT-SoVITS is the best if you can put in the effort. I'm still using v2 since the output is so good. Its also the best multilingual one in my testing on Japanese inputs.

can it support more languages rather than only English, Chinese, Japanese, Korean?
Post reply on HN