Live data from Hacker News

Eleven v3

elevenlabs.io

151–160 of 168 posts

Re: Eleven v3

#151

Earlier quoted context omitted.

> The result with OpenAI feels much less predictable and of lower production quality than ElevenLabs Thank you Ian! Credit to our research team for making this possible For the prosidy, if you choose an expressive voice the prosidy should be larger

Ninjaing in to ask: is v3 on the roadmap for your voice agents? The quality increase is huge.

Yep, low latency models are on the way.

Re: Eleven v3

#153

Earlier quoted context omitted.

ElevenLabs v2's accented voices are still much stronger than any of its competition. And I've tried it with Arabic, French, Hindi and English.

Can it do a proper Singaporean or Hongkongese accent?

Haven't tried it, but it does an Arabic-accented English somewhat okayishly.

Re: Eleven v3

#154

Sounds absolutely amazing, like 99% indistinguishable from real professional voice actors to me. I couldn't find any pricing though. Anyone know what they charge for it?

But it's not an actual person. It's an "AI". Do you want a future where you don't hear actual people anymore? I want to listen to music, audiobooks, poetry, novels, plays, with actual humans talking, that's the whole fucking point.

I feel like you're conflating the act of creation (writing a book) versus the act of performance (narrating the book). For the former I agree with you, but for the latter? Shrug.

Personally I have hundreds of old texts that simply do not have an audio book equivalent and using realistic sounding TTS has been perfectly adequate.

Re: Eleven v3

#156

Earlier quoted context omitted.

Yeah it's irritating enough when humans do it, it's so transparently insincere. Just help me with my problem. I guess I am just old now but I hate talking to computers, I never use Siri or any other voice interfaces, and I don't want computers talking to me as if they are human. Maybe if it were like Star Trek and the computer just said "Working..." and then gave me the answer it would be tolerable. Just please cut o…

I agree it seems transparently insincere yes, but the reason it’s done is because it works on some people who either don’t detect it or need it as politeness norms and the ones who see it as insincere just ignore it and move on. Thus net, you win by doing this because it rarely if ever costs you and thus you only have upside.

> The ones who see it as insincere just ignore it and move on.

Except I "just move on" to another product.

The only person I know who doesn't find this pretension annoying is my 90 year-old mother. I don't have time to waste on any company that wastes my time with pointless cut-and-paste babble. And any company actually intentionally catering to my 90 year-old mother as a primary target customer is clearly signaling they aren't for me.

A decade from now such blatant condescension from an AI will be a trope: "OMG, that's so mid-2020s AI it's painful."

Re: Eleven v3

#157

I've been using OpenAI's new models a lot lately ( https://www.openai.fm/ )... separating instructions from the spoken word is an interesting choice, and I'm assuming also has a lot to do with OpenAI/GPT using "instructions" across their products, and maybe they are just more comfortable and familiar generating the data and do the training for that style. Separate instructions is a bit awkward, but does allow mixing…

> But in the end OpenAI's biggest feature is that it's 10x cheaper and completely pay-as-you-go. (Why are all these TTS services doing subscriptions on top of limits and credits? Blech!) Is it so, after all the LLM and overheads have been considered? Elevenlabs conversational agents are priced at 0.08 per minute at the highest tier. How much is the comparable at Open AI? I did a rough estimate and found it was higher…

It's confusing, but if I look closer then 10x is an exaggeration, it's more like 5x...

https://elevenlabs.io/pricing

Creator tier (lowest tier that's full service) is $22/mo for 250 minutes, $0.08/minute. Then it's $0.15/1000 characters. (So many different fucking units! And these prices are actually "credits" translated to other units; I fucking hate funny-money "credits")

https://platform.openai.com/docs/pricing#transcription-and-s...

Estimated $0.015/minute (actually priced based on tokens; yet more weird units!)

The non-instruction models are $0.015/1000 characters.

It starts getting more competitive when you are at the highest tier at ElevenLabs ($1320/month), but because of their pricing structure I'm not going to invest the time in finding out if it's worth it.

Re: Eleven v3

#158
post #98

Earlier quoted context omitted.

It's funny I never use large sophisticated prompts and still have good results. Something like: > Always be concise and trust that I will understand what you say on the first try. No fluff in your answers, speak directly to the point. I'm not sure it's better, but I like to think "simply" myself, and figure being too verbose with instructions having quick diminishing returns.

What's more likely to be a problem, is the request to be concise. For some reason, this still seems to not be widely known among even technical users: token generation is where the computation/"thinking" in LLMs happen! By forcing it to keep its answers short, you're starving the model for compute, making each token do more work. There's a small, fixed amount of "thinking" LLM can do per token, so the more you squeez…

If that's correct then it's a significant problem with LLMs that needs to be addressed. Would it work to have the agent keep the talky, verbose answer to itself and only return to a finally summary to the user?

Re: Eleven v3

#159

Earlier quoted context omitted.

I agree it seems transparently insincere yes, but the reason it’s done is because it works on some people who either don’t detect it or need it as politeness norms and the ones who see it as insincere just ignore it and move on. Thus net, you win by doing this because it rarely if ever costs you and thus you only have upside.

> The ones who see it as insincere just ignore it and move on. Except I "just move on" to another product. The only person I know who doesn't find this pretension annoying is my 90 year-old mother. I don't have time to waste on any company that wastes my time with pointless cut-and-paste babble. And any company actually intentionally catering to my 90 year-old mother as a primary target customer is clearly signaling…

It will be a trope eventually. But like I said, the cost benefit analysis puts it generally in the benefit camp. And if every next product also does this, are you actually going to not use the product? In most cases I think people put up with this & just minimize the interaction that leads to this (another benefit for the support team wording things this way since they have to field fewer support requests)

Re: Eleven v3

#160

Earlier quoted context omitted.

> But in the end OpenAI's biggest feature is that it's 10x cheaper and completely pay-as-you-go. (Why are all these TTS services doing subscriptions on top of limits and credits? Blech!) Is it so, after all the LLM and overheads have been considered? Elevenlabs conversational agents are priced at 0.08 per minute at the highest tier. How much is the comparable at Open AI? I did a rough estimate and found it was higher…

It's confusing, but if I look closer then 10x is an exaggeration, it's more like 5x... https://elevenlabs.io/pricing Creator tier (lowest tier that's full service) is $22/mo for 250 minutes, $0.08/minute. Then it's $0.15/1000 characters. (So many different fucking units! And these prices are actually "credits" translated to other units; I fucking hate funny-money "credits") https://platform.openai.com/docs/pricing#tr…

> It starts getting more competitive when you are at the highest tier at ElevenLabs ($1320/month), but because of their pricing structure I'm not going to invest the time in finding out if it's worth it.

They do have a grant programme through, which gives 3 months free of the largest tier.

https://elevenlabs.io/startup-grants

Post reply on HN