Earlier quoted context omitted.
If you compare with e.g. Deepseek and other hosters, you'll find that OpenAI is actually almost certainly charging very high margins (Deepseek has an 80% profit margin and they're 10x cheaper than openai). The training/R&D might make OpenAI burn VC cash, but this isn't comparable with companies like WeWork whose products actively burn cash
They said themselves that even inference is losing them money tho, or did I get that wrong?
OpenAI Audio Models
281–290 of 317 posts
Re: OpenAI Audio Models
#282Hey, I'm Jeff and I was PM for these models at OpenAI. Today we launched three new state-of-the-art audio models. Two speech-to-text models—outperforming Whisper. A new TTS model—you can instruct it how to speak (try it on openai.fm!). And our Agents SDK now supports audio, making it easy to turn text agents into voice agents. We think you'll really like these models. Let me know if you have any questions here!
Another thing I noticed is whisper did a better job of transcribing when I removed a lot of the silences in the audio.
Re: OpenAI Audio Models
#283Both the text-to-speech and the speech-to-text models launched here suffer from reliability issues due to combining instructions and data in the same stream of tokens. I'm not yet sure how much of a problem this is for real-world applications. I wrote a few notes on this here: https://simonwillison.net/2025/Mar/20/new-openai-audio-model...
Re: OpenAI Audio Models
#284Earlier quoted context omitted.
Did you try Kokoro? You can self host that. https://huggingface.co/spaces/hexgrad/Kokoro-TTS
Thanks! But I get the impression that with Kokoro, a strong CPU still requires about two seconds to generate one sentence, which is too much of a delay for a TTS voice in an AAC app. I'd rather accept a little compromise regarding the voice and intonation quality, as long as the TTS system doesn't frequently garble words. The AAC app is used on tablet PCs running from battery, so the lower the CPU usage and energy dr…
Re: OpenAI Audio Models
#285Large text-to-speech and speech-to-text models have been greatly improving recently. But I wish there were an offline , on-device, multilingual text-to-speech solution with good voices for a standard PC — one that doesn't require a GPU, tons of RAM, or max out the CPU. In my research, I didn't find anything that fits the bill. People often mention Tortoise TTS, but I think it garbles words too often. The only plug-in…
Re: OpenAI Audio Models
#286Earlier quoted context omitted.
not open source at this time. unfortunately they're much to large to run on normal consumer hardware
Is that the reason you're not open sourcing them? Wouldn't it still make sense to provide it for enthusiasts?
Re: OpenAI Audio Models
#287Re: OpenAI Audio Models
#288The site just crashes with service workers disabled. First time I ran into a problem with that setting, which I set over two years ago.
oh doh. thanks ... we just pushed a fix for the crash. Unfortunately our currently implementation needs service works for streaming audio, so the "fix" was to disable the feature if the worker isn't available
Streaming audio is a new one to me, I wonder if the same could be achieved with web workers instead. Or at least similar use cases like video calls work fine for me without service workers. See e.g. https://github.com/scottstensland/web-audio-workers-sockets
Re: OpenAI Audio Models
#289This is astonishing. I can type anything I want into the "vibe" box and it does it for the given text. Accents, attitudes, personality types... I'm amazed. The level of intelligent "prosody" here -- the rhythm and intonation, the pauses and personality -- I wasn't expecting anything like this so soon. This is truly remarkable. It understands both the text and the prompt for how the speaker should sound. Like, we're g…
Vibe:
Voice Affect: A Primal Scream from the top of your lungs!
Tone: LOUD. A RAW SCREAM
Emotion: Intense primal rage.
Pronunciation: Draw out the last word until you are out of breath.
Script:
EVERY THING WAS SAD!
Re: OpenAI Audio Models
#290This is astonishing. I can type anything I want into the "vibe" box and it does it for the given text. Accents, attitudes, personality types... I'm amazed. The level of intelligent "prosody" here -- the rhythm and intonation, the pauses and personality -- I wasn't expecting anything like this so soon. This is truly remarkable. It understands both the text and the prompt for how the speaker should sound. Like, we're g…
I've been trying to get it to scream with some humorous results: Vibe: Voice Affect: A Primal Scream from the top of your lungs! Tone: LOUD. A RAW SCREAM Emotion: Intense primal rage. Pronunciation: Draw out the last word until you are out of breath. Script: EVERY THING WAS SAD!