Live data from Hacker News

This Voice Doesn't Exist – Generative Voice AI

blog.elevenlabs.io

121–130 of 279 posts

Re: This Voice Doesn't Exist – Generative Voice AI

#121
post #111

Okay can I ask a question that has been bothering me for a long time? Why do seemingly all these text-to-speech programs attempt to produce spoken voice based solely on raw text ? Why don't they consume a MIDI-like text-markup language where you can write phonetic pronunciations along with markup about the emotion, volume, speed, etc.? I feel like this is a huge unnecessary roadblock holding back this kind of technol…

[deleted]

Re: This Voice Doesn't Exist – Generative Voice AI

#122
post #111

Okay can I ask a question that has been bothering me for a long time? Why do seemingly all these text-to-speech programs attempt to produce spoken voice based solely on raw text ? Why don't they consume a MIDI-like text-markup language where you can write phonetic pronunciations along with markup about the emotion, volume, speed, etc.? I feel like this is a huge unnecessary roadblock holding back this kind of technol…

Just my 2 cents, but it seems to me that too little focus in the tech world has been spent on understannding what speech is. Tonality, mood, facial expression and body language all is ignored or people pretend like there are no such thing. I believe this is broadly true in western society by now- people went digital but do not yet realize why communication went to hell in the last decade.

Re: This Voice Doesn't Exist – Generative Voice AI

#123
post #109

Earlier quoted context omitted.

> what are the alternatives that I can run locally ...you will be disappointed by the answers to that question for the foreseeable future.

I'm the opposite of disappointed. The amount of public pretrained models that have been popping up recently is crazy.

[deleted]

Re: This Voice Doesn't Exist – Generative Voice AI

#124
post #109

Earlier quoted context omitted.

> what are the alternatives that I can run locally ...you will be disappointed by the answers to that question for the foreseeable future.

I'm the opposite of disappointed. The amount of public pretrained models that have been popping up recently is crazy.

Same model with random tweaks applied.

Just because there is a new toy doesn’t mean capitalism gave up.

Re: This Voice Doesn't Exist – Generative Voice AI

#125
The text to speech function at the top of the article is the actual product but they are not going the extra mile and record it again for the other speed multiplier like x.7 or x2.0. You can clearly hear the mp3 struggling, especially at 0.7 speed.

It would have been interesting how they perform in comparison. The fact you are able to adjust the voices is even one of their selling points. I really wonder why they haven't done that

Re: This Voice Doesn't Exist – Generative Voice AI

#126
post #111

Okay can I ask a question that has been bothering me for a long time? Why do seemingly all these text-to-speech programs attempt to produce spoken voice based solely on raw text ? Why don't they consume a MIDI-like text-markup language where you can write phonetic pronunciations along with markup about the emotion, volume, speed, etc.? I feel like this is a huge unnecessary roadblock holding back this kind of technol…

I don’t know how many of the solutions offer this, but there is a markup language for TTS: https://en.wikipedia.org/wiki/Speech_Synthesis_Markup_Langua... Amazon Polly, (which seems kind of ancient with all these new solutions showing up) has supported SSML for some time. AWS Polly SSML docs: https://docs.aws.amazon.com/polly/latest/dg/ssml.html

In practice they are next to useless, the expressions are not very...expressive (just try it in the AWS editor). I suspect a LLM would be able to infer the context or we can use prompt engineering to generate the appropriate tokens encoding emotions for the intermediate neural codecs directly (Mel spectrograms are so passé now post Vall-E).

Re: This Voice Doesn't Exist – Generative Voice AI

#127

What are the odds of this kind of thing being open source so I can use it at home. So far, most of the "good" text-to-speech systems are all commercial services https://aws.amazon.com/polly/ https://cloud.google.com/text-to-speech https://azure.microsoft.com/en-us/products/cognitive-service... And now one is also a service. I tried using tortoise-tts on my M1. Generating a 7 minute speech took 3 days and, while bette…

That’s because you’re running tortoise on a CPU. It does about a sentence a minute on my 3090 gpu. It’s also quite good if you pick “high quality” and train it with 10 sec clips at the framerate and bitrate it asks for.

Re: This Voice Doesn't Exist – Generative Voice AI

#128

Earlier quoted context omitted.

I want to agree, but I searched on their website and found their narration service with 2 full book examples. I listened to the first one for a while and it's the first time an Ai narrator was good enough to keep me listening: https://www.audiostory.ai/2065785/11707800-alice-s-adventure...

It's noticable worse than the examples in the blog post. I mean, it's good enough for listening, but no better than the competition.

[dead]

Re: This Voice Doesn't Exist – Generative Voice AI

#130

I'm listening to an audiobook whose reader is not as good as some of these voices. At one level, I'm impressed but at an another I'm sadden since we are heading towards uncharted territory. We are looking at a future where we'll have content, video,audio, and text by the truckload. More does not mean better. It just means more blah stuff. I don't think that's the future I'm looking forward to live in.

[dead]
Post reply on HN