Live data from Hacker News

OpenAI Audio Models

openai.fm

151–160 of 317 posts

Re: OpenAI Audio Models

#151
post #85

It doesn't seem clear, but can the model do correct emphesis? On things like single words: I did not steal that horse Is the trivial example of something where intonation of the single word is what matters. More importantly if you are reading something, as a human, you change the intonation, audiolevel, and speed.

> I did not steal that horse

> Is the trivial example of something where intonation of the single word is what matters.

My go-to for an example of this is "I didn't say she stole my money".

Changing which word is emphasized completely changes the meaning of the sentence.

Re: OpenAI Audio Models

#152

If I'm reading the pricing correctly, these models are SIGNIFICANTLY cheaper than ElevenLabs. https://platform.openai.com/docs/pricing If these are the "gpt-4o-mini-tts" models, and if the pricing estimate of "$0.015 per minute" of audio is correct, then these prices 85% cheaper than those of ElevenLabs. https://elevenlabs.io/pricing With ElevenLabs, if I choose their most cost-effectuve "Business" plan for $1100 per…

Sesame is free and pretty good and you can run it yourself.

They released a crippled model: https://github.com/SesameAILabs/csm/issues/63

Re: OpenAI Audio Models

#153

Earlier quoted context omitted.

Yes, from our terms: "Don’t build tools that may be inappropriate for minors, including: Sexually explicit or suggestive content. This does not include content created for scientific or educational purposes." https://openai.com/policies/usage-policies/

[flagged]

The general consensus of AI overlords is that humans are minors.

Re: OpenAI Audio Models

#154

Earlier quoted context omitted.

Does whispering work? I could not get it to work when I tried it

Should do! here's an example https://www.openai.fm/#4a5a82db-faea-4f80-813c-3131902c2458

It seems to start out strong, but then starts loudly talking by the end, do you know why it loses focus?

edit: I actually got it to stay whispering by also putting (soft whispering voice) before the second paragraph

Re: OpenAI Audio Models

#156
post #141

Large text-to-speech and speech-to-text models have been greatly improving recently. But I wish there were an offline , on-device, multilingual text-to-speech solution with good voices for a standard PC — one that doesn't require a GPU, tons of RAM, or max out the CPU. In my research, I didn't find anything that fits the bill. People often mention Tortoise TTS, but I think it garbles words too often. The only plug-in…

Look into https://superwhisper.com and their local models. Pretty decent.

Thank you, but they say "Offline models only run really well on Apple Silicon macs."

Re: OpenAI Audio Models

#157
post #98
post #79

Earlier quoted context omitted.

That's a really cool page thanks. Does it have stats for other languages? In my experience the OpenAI TTS APIs were really bad, messing up all the time in foreign languages. Practically unusable for my use case. You'd have to use the gpt-4o-audio-preview to get anything close to passable, but it was expensive. Which is why I'm using Google TTS which is very fast, high quality, and provides first class support for alm…

Interesting for me Open TTS for Polish was better than Google TTS (but they have few options) - which one did you used? WaveNet? Sadly haven't seen quality evaluation for TTS for foreign languages

Depends on what's available for the language, but yea Wavenet and Neural2. With OpenAI TTS I'd often get weird bugs where the first API call comes back all garbled, but the second API call comes back fine. Wasting money. On top of that more expensive and higher latency. I'm interested to try out this new one.

Re: OpenAI Audio Models

#158
post #146
post #141

Large text-to-speech and speech-to-text models have been greatly improving recently. But I wish there were an offline , on-device, multilingual text-to-speech solution with good voices for a standard PC — one that doesn't require a GPU, tons of RAM, or max out the CPU. In my research, I didn't find anything that fits the bill. People often mention Tortoise TTS, but I think it garbles words too often. The only plug-in…

May I introduce to you https://huggingface.co/canopylabs/orpheus-3b-0.1-ft (no affiliation) it's English only afaics.

The sample sounds impressive, but based on their claim -- 'Streaming inference is faster than playback even on an A100 40GB for the 3 billion parameter model' -- I don't think this could run on a standard laptop.

Re: OpenAI Audio Models

#160
is it just me or are these voices clearly AI generated? They've obviously been improving at a steady rate but if I saw a YouTube video that had this voice, I'd instantly stop watching it
Post reply on HN