Earlier quoted context omitted.
> The result with OpenAI feels much less predictable and of lower production quality than ElevenLabs Thank you Ian! Credit to our research team for making this possible For the prosidy, if you choose an expressive voice the prosidy should be larger
Ninjaing in to ask: is v3 on the roadmap for your voice agents? The quality increase is huge.
Eleven v3
151–160 of 168 posts
Re: Eleven v3
#152Re: Eleven v3
#153Earlier quoted context omitted.
ElevenLabs v2's accented voices are still much stronger than any of its competition. And I've tried it with Arabic, French, Hindi and English.
Can it do a proper Singaporean or Hongkongese accent?
Re: Eleven v3
#154Sounds absolutely amazing, like 99% indistinguishable from real professional voice actors to me. I couldn't find any pricing though. Anyone know what they charge for it?
But it's not an actual person. It's an "AI". Do you want a future where you don't hear actual people anymore? I want to listen to music, audiobooks, poetry, novels, plays, with actual humans talking, that's the whole fucking point.
Personally I have hundreds of old texts that simply do not have an audio book equivalent and using realistic sounding TTS has been perfectly adequate.
Re: Eleven v3
#155The actual title of the release: Eleven v3 -- The most expensive Text to Speech model
Re: Eleven v3
#156Earlier quoted context omitted.
Yeah it's irritating enough when humans do it, it's so transparently insincere. Just help me with my problem. I guess I am just old now but I hate talking to computers, I never use Siri or any other voice interfaces, and I don't want computers talking to me as if they are human. Maybe if it were like Star Trek and the computer just said "Working..." and then gave me the answer it would be tolerable. Just please cut o…
I agree it seems transparently insincere yes, but the reason it’s done is because it works on some people who either don’t detect it or need it as politeness norms and the ones who see it as insincere just ignore it and move on. Thus net, you win by doing this because it rarely if ever costs you and thus you only have upside.
Except I "just move on" to another product.
The only person I know who doesn't find this pretension annoying is my 90 year-old mother. I don't have time to waste on any company that wastes my time with pointless cut-and-paste babble. And any company actually intentionally catering to my 90 year-old mother as a primary target customer is clearly signaling they aren't for me.
A decade from now such blatant condescension from an AI will be a trope: "OMG, that's so mid-2020s AI it's painful."
Re: Eleven v3
#157I've been using OpenAI's new models a lot lately ( https://www.openai.fm/ )... separating instructions from the spoken word is an interesting choice, and I'm assuming also has a lot to do with OpenAI/GPT using "instructions" across their products, and maybe they are just more comfortable and familiar generating the data and do the training for that style. Separate instructions is a bit awkward, but does allow mixing…
> But in the end OpenAI's biggest feature is that it's 10x cheaper and completely pay-as-you-go. (Why are all these TTS services doing subscriptions on top of limits and credits? Blech!) Is it so, after all the LLM and overheads have been considered? Elevenlabs conversational agents are priced at 0.08 per minute at the highest tier. How much is the comparable at Open AI? I did a rough estimate and found it was higher…
Creator tier (lowest tier that's full service) is $22/mo for 250 minutes, $0.08/minute. Then it's $0.15/1000 characters. (So many different fucking units! And these prices are actually "credits" translated to other units; I fucking hate funny-money "credits")
https://platform.openai.com/docs/pricing#transcription-and-s...
Estimated $0.015/minute (actually priced based on tokens; yet more weird units!)
The non-instruction models are $0.015/1000 characters.
It starts getting more competitive when you are at the highest tier at ElevenLabs ($1320/month), but because of their pricing structure I'm not going to invest the time in finding out if it's worth it.
Re: Eleven v3
#158Earlier quoted context omitted.
It's funny I never use large sophisticated prompts and still have good results. Something like: > Always be concise and trust that I will understand what you say on the first try. No fluff in your answers, speak directly to the point. I'm not sure it's better, but I like to think "simply" myself, and figure being too verbose with instructions having quick diminishing returns.
What's more likely to be a problem, is the request to be concise. For some reason, this still seems to not be widely known among even technical users: token generation is where the computation/"thinking" in LLMs happen! By forcing it to keep its answers short, you're starving the model for compute, making each token do more work. There's a small, fixed amount of "thinking" LLM can do per token, so the more you squeez…
Re: Eleven v3
#159Earlier quoted context omitted.
I agree it seems transparently insincere yes, but the reason it’s done is because it works on some people who either don’t detect it or need it as politeness norms and the ones who see it as insincere just ignore it and move on. Thus net, you win by doing this because it rarely if ever costs you and thus you only have upside.
> The ones who see it as insincere just ignore it and move on. Except I "just move on" to another product. The only person I know who doesn't find this pretension annoying is my 90 year-old mother. I don't have time to waste on any company that wastes my time with pointless cut-and-paste babble. And any company actually intentionally catering to my 90 year-old mother as a primary target customer is clearly signaling…
Re: Eleven v3
#160Earlier quoted context omitted.
> But in the end OpenAI's biggest feature is that it's 10x cheaper and completely pay-as-you-go. (Why are all these TTS services doing subscriptions on top of limits and credits? Blech!) Is it so, after all the LLM and overheads have been considered? Elevenlabs conversational agents are priced at 0.08 per minute at the highest tier. How much is the comparable at Open AI? I did a rough estimate and found it was higher…
It's confusing, but if I look closer then 10x is an exaggeration, it's more like 5x... https://elevenlabs.io/pricing Creator tier (lowest tier that's full service) is $22/mo for 250 minutes, $0.08/minute. Then it's $0.15/1000 characters. (So many different fucking units! And these prices are actually "credits" translated to other units; I fucking hate funny-money "credits") https://platform.openai.com/docs/pricing#tr…
They do have a grant programme through, which gives 3 months free of the largest tier.