Live data from Hacker News

OpenAI Audio Models

openai.fm

271–280 of 317 posts

Re: OpenAI Audio Models

#271
post #233

Earlier quoted context omitted.

> Not sure how that's possible Download bunch of movies Scarlet Johansen been in, segment into audio clips where she talks and train the model :)

Is it actually her? I didn't think it was, but maybe.

Unless there is some leak from OpenAI, I'm not sure we'll ever have it confirmed yes or no. But my brain thought it was Johansen from the first few seconds I heard the voice and I don't seem to be alone with that reaction. The fact that they removed the voice also speaks to it to have been trained on her voice.

Listening to it again today with fresher ears (the original OpenAI Sky, not the clones elsewhere), I still hear Johansen as the underlying voice actor for it, but maybe there is some subconscious bias I'm unable to bypass.

Re: OpenAI Audio Models

#272

Still seems like Elevenlabs is crushing them on realtime audio, or does this change things?

I'm also curious about this for longform content. Will this be competitive for something like creating an audiobook?

For $0.015 a minute it has to be.

The books I am listening to now wouldn't even be $10. Any future price drops then will really make this a no-brainer.

The Elevenlabs pricing to me makes it completely useless for audiobooks that I just want to listen to for my personal enjoyment.

Re: OpenAI Audio Models

#273

Earlier quoted context omitted.

It's cheap because everything OpenAI does is subsidized by investors' money. Until that stupid money flows all good! Then either they'll go the way of WeWork, or enshittification will happen to make it possible for them to make the books work. I don't see any other option. Unless Softbank decides it has some 150 Billion to squander on buying them off. There's a lot of head-in-the-sand behavior going on around OpenAI…

If you compare with e.g. Deepseek and other hosters, you'll find that OpenAI is actually almost certainly charging very high margins (Deepseek has an 80% profit margin and they're 10x cheaper than openai). The training/R&D might make OpenAI burn VC cash, but this isn't comparable with companies like WeWork whose products actively burn cash

They said themselves that even inference is losing them money tho, or did I get that wrong?

Re: OpenAI Audio Models

#274

Earlier quoted context omitted.

yeah it's almost like an uncanny valley where it sometimes feels like the voice trying to be an actor and play a character or something

I mean that's literally the service it's providing. If you asked humans to do the same thing it would sound equally forced. All acting sounds cringe out of context.

IDK, i've done my fair share of amateur acting and at least to me (english is not my first language) there's something more uncanny than just the typical "say this wihtout knowing much of the context)

Re: OpenAI Audio Models

#275
post #271

Earlier quoted context omitted.

Is it actually her? I didn't think it was, but maybe.

Unless there is some leak from OpenAI, I'm not sure we'll ever have it confirmed yes or no. But my brain thought it was Johansen from the first few seconds I heard the voice and I don't seem to be alone with that reaction. The fact that they removed the voice also speaks to it to have been trained on her voice. Listening to it again today with fresher ears (the original OpenAI Sky, not the clones elsewhere), I still…

Hmm, I never thought it was her, her voice is much more raspy, whereas Sky is a bit lighter. I can hear the similarity, I just don't think they sound exactly alike.

As you say, I'm not sure we'll ever know, although the Sky voice from Kokoro is spot on the Sky voice from OpenAI, so maybe someone from Kokoro knows how they got it.

Re: OpenAI Audio Models

#277
A bit off-topic but I'm so glad to see skeuomorphic UI make a comeback!

Check out the toggle switch in the upper right corner! I hope more designers will follow this example.

Re: OpenAI Audio Models

#278
post #223

Hey, I'm Jeff and I was PM for these models at OpenAI. Today we launched three new state-of-the-art audio models. Two speech-to-text models—outperforming Whisper. A new TTS model—you can instruct it how to speak (try it on openai.fm!). And our Agents SDK now supports audio, making it easy to turn text agents into voice agents. We think you'll really like these models. Let me know if you have any questions here!

Hi Jeff. This is awesome. Any plans to add word timestamps to the new speech-to-text models, though? > Other parameters, such as timestamp_granularities, require verbose_json output and are therefore only available when using whisper-1. Word timestamps are insanely useful for large calls with interruptions (e.g. multi-party debate/Twitter spaces), allowing transcript lines to be further split post-transcription on se…

You need speaker attribution, right?

Re: OpenAI Audio Models

#279

This is astonishing. I can type anything I want into the "vibe" box and it does it for the given text. Accents, attitudes, personality types... I'm amazed. The level of intelligent "prosody" here -- the rhythm and intonation, the pauses and personality -- I wasn't expecting anything like this so soon. This is truly remarkable. It understands both the text and the prompt for how the speaker should sound. Like, we're g…

Didn’t look closely, but is there a way to clone a voice from a few seconds of recording and then feed the sample to generate the text in the same voice?

Re: OpenAI Audio Models

#280

Earlier quoted context omitted.

If you compare with e.g. Deepseek and other hosters, you'll find that OpenAI is actually almost certainly charging very high margins (Deepseek has an 80% profit margin and they're 10x cheaper than openai). The training/R&D might make OpenAI burn VC cash, but this isn't comparable with companies like WeWork whose products actively burn cash

They said themselves that even inference is losing them money tho, or did I get that wrong?

On their subscriptions, specifically the pro subscription, because it's a flatrate to their most expensive model. The API prices are all much more expensive. It's unclear whether they're losing money on the normal subscriptions, but if so, probably not by much. Though it's definitely closer to what you described, subsidizing it to gain 'mindshare' or whatever.
Post reply on HN