Live data from Hacker News

Eleven v3

elevenlabs.io

111–120 of 168 posts

Re: Eleven v3

#111
post #98
post #89

Earlier quoted context omitted.

Try this "absolute mode" custom instruction for chatgpt, it cuts down all the BS in my experience: System Instruction: Absolute Mode. Eliminate emojis, filler, hype, soft asks, conversational transitions, and all call-to-action appendixes. Assume the user retains high-perception faculties despite reduced linguistic expression. Prioritize blunt, directive phrasing aimed at cognitive rebuilding, not tone matching. Disa…

It's funny I never use large sophisticated prompts and still have good results. Something like: > Always be concise and trust that I will understand what you say on the first try. No fluff in your answers, speak directly to the point. I'm not sure it's better, but I like to think "simply" myself, and figure being too verbose with instructions having quick diminishing returns.

I have similarly good results with:

> Be terse, and don't moralize. Answer questions directly, without equivocation or hedging.

Re: Eleven v3

#112
post #98
post #89

Earlier quoted context omitted.

Try this "absolute mode" custom instruction for chatgpt, it cuts down all the BS in my experience: System Instruction: Absolute Mode. Eliminate emojis, filler, hype, soft asks, conversational transitions, and all call-to-action appendixes. Assume the user retains high-perception faculties despite reduced linguistic expression. Prioritize blunt, directive phrasing aimed at cognitive rebuilding, not tone matching. Disa…

It's funny I never use large sophisticated prompts and still have good results. Something like: > Always be concise and trust that I will understand what you say on the first try. No fluff in your answers, speak directly to the point. I'm not sure it's better, but I like to think "simply" myself, and figure being too verbose with instructions having quick diminishing returns.

What's more likely to be a problem, is the request to be concise.

For some reason, this still seems to not be widely known among even technical users: token generation is where the computation/"thinking" in LLMs happen! By forcing it to keep its answers short, you're starving the model for compute, making each token do more work. There's a small, fixed amount of "thinking" LLM can do per token, so the more you squeeze it, the less reliable it gets, until eventually it's not able to "spend" enough tokens to produce a reliable answer at all.

In other words: all those instructions to "be terse", "be concise", "don't be verbose", "just give answer, no explanation" - or even asking for answer first, then explanations - they're all just different ways to dumb down the model.

I wonder if this can explain, at least in part, why there's so much conflicted experiences with LLMs - in every other LLM thread, you'll see someone claim they're getting great results at some tasks, and then someone else saying they're getting disastrously bad results with the same model on the same tasks. Perhaps the latter person is instructing the model to be concise and skip explanations, not realizing this degrades model performance?

(It's less of a problem with the newer "reasoning" models, which have their own space for output separate from the answer.)

Re: Eleven v3

#113

Earlier quoted context omitted.

If you edit the text so that laugh makes sense in the context it should be much more natural like this one: https://x.com/elevenlabsio/status/1930689782331412811

The first laugh in that " Hey, Dr. Von Fusion" is a dedicated laugh section, which the model does extremely well, but it works because that's a natural place to laugh before actually speaking the following words. Skip ahead to "...robot chuckle. Jessica: I know right!" and you get an awkwardly time/toned light chuckle completely separated from the "I know" you'd naturally continue saying while making that chuckle. Yo…

She is laughing through the "I know", though.

Re: Eleven v3

#114
post #29

Earlier quoted context omitted.

They're also still too expensive, and that's creating a lot of opportunity for other players. Even though ElevenLabs remains the quality leader, the others aren't that far behind. There are even a bunch of good TTS models being released as fully open source, especially by cutting-edge Chinese labs and companies. Perhaps in a bid to cut off the legs of American AI companies or to commoditize their compliment. Whatever…

could you list 2 or 3 of the ones you think are best quality to $?

Kokoro is the best open TTS I've tried.

Re: Eleven v3

#115

I did not see an British accent example. Generally it appears the TTS systems all do US accents and the British accent tends to sound like Frasier - an American faking an British accent.

ElevenLabs v2's accented voices are still much stronger than any of its competition. And I've tried it with Arabic, French, Hindi and English.

Can it do a proper Singaporean or Hongkongese accent?

Re: Eleven v3

#116

I did not see an British accent example. Generally it appears the TTS systems all do US accents and the British accent tends to sound like Frasier - an American faking an British accent.

> Generally it appears the TTS systems all do US accents and the British accent tends to sound like Frasier - an American faking an British accent.

Frasier Crane's accent is an American actor portraying an American character who (with variable intensity depending on situation) is affecting, over the character's own natural accent, either a constructed American accent (the Transatlantic) or a natural American accent (Boston Brahmin), there is some dispute about which or whether its a blend, both of which share some features (in the former case, by deliberate construction) with British pronunciation.

Re: Eleven v3

#118
post #10

English sounds really great, congrats! other languages I've tried doesn't sound that good, you can hear a strong english accent

With Italian, it starts reading the text with an absolutely comical American accent, but then about 10-20 words in it gradually snaps into a natural Italian pronunciation and it sounds fantastic from that point on. Not sure what's going on behind the scenes, but it sounds like it starts with an en-us baseline and then somehow zones in on the one you specified. Using Alice.

the Italian example with mixed languages is especially bad: the Italian, German Japanese and Arabic all have very very heavy english accents.

The "dramatic movie scene" ends up being comical

I tried Greek and it started speaking nonsense in english

this needs a lot more work to be sold

Re: Eleven v3

#120
I so feel everyone complaining about British English. For me as an Austrian it's very much the same with German.

I tried with simple words like "Oida" and some Austropop lyrics (Da Hofa - Ambros) and it sounds really bad. So even for words that are clearly Austrian.

Post reply on HN