Live data from Hacker News

PlayAI's new Dialog model achieves 3:1 preference in human evals

play.ht

21–30 of 59 posts

Re: PlayAI's new Dialog model achieves 3:1 preference in human evals

#21

For some reason, most of these (and other narration AIs) sound like someone reading off a teleprompter, rather than natural speaking voices. I'm not sure what exactly it is, but I'm left feeling like the speaker isn't really sure of what the next words are, and the stresses between the words are all over the place. It's like the emphasis over a sentence doesn't really match how humans sound.

You have to shape the voices in the tools, if you just spit them out they're junk but if you take the time to shape the voice a bit it gets better quickly, this is a cheap 11labs voice with 30 seconds spent on some basic shaping: https://s.h4x.club/bLuNlJWx

Still a bit teleprompter-ish but there are tools to go in and adjust pace and style throughout and you probably hear a lot of stuff with people not using those creative features. 11labs might very well be one of the best bits of software I've used, it's a great deal of fun to play with and if you're willing to spend the time the results are superb - I don't even have a use case, I just like making them because they're fun to listen to, ha!

Re: PlayAI's new Dialog model achieves 3:1 preference in human evals

#22
post #18

I've been messing with the open source side of audio generation, and expressiveness still takes work but it's getting there. Roughly summarized my findings are: - zero shot voice cloning isn't there yet - gpt-sovits is the best at non-word vocalizations, but the overall quality is bad when just using zero shot, finetuning helps - F5 and fish-speech are both good as well - xtts for me has had the best stability (i can…

Xtts is non commercial use only though

I think XTTS is MPL now since Coqui folded, but I am not a lawyer and I am not using this for anything commercial so I haven't looked closely.

Re: PlayAI's new Dialog model achieves 3:1 preference in human evals

#23
post #17
post #15

I tried it with a paragraph of English taken from a formal speech, and it sounded quite good. I would not have been able to distinguish it from a skilled human narrator. But then I tried a paragraph of Japanese text, also from a formal speech, with the language set to Japanese and the narrator set to Yumiko Narrative. The result was a weird mixture of Korean, Chinese, and Japanese readings for the kanji and kana, all…

Yeah there are no good options for Japanese yet (except maybe in Japan but I haven't heard of good Ai models for speech locally)

Anecdotally gpt-sovits is quite good at japanese, I can't evaluate first hand as my japanese is trash.

Re: PlayAI's new Dialog model achieves 3:1 preference in human evals

#24
post #15

I tried it with a paragraph of English taken from a formal speech, and it sounded quite good. I would not have been able to distinguish it from a skilled human narrator. But then I tried a paragraph of Japanese text, also from a formal speech, with the language set to Japanese and the narrator set to Yumiko Narrative. The result was a weird mixture of Korean, Chinese, and Japanese readings for the kanji and kana, all…

So kind of unrelated, but the reading/singing of arbitrary custom lyrics on suno.com's v4 model has blown me away.

Re: PlayAI's new Dialog model achieves 3:1 preference in human evals

#25
post #21

For some reason, most of these (and other narration AIs) sound like someone reading off a teleprompter, rather than natural speaking voices. I'm not sure what exactly it is, but I'm left feeling like the speaker isn't really sure of what the next words are, and the stresses between the words are all over the place. It's like the emphasis over a sentence doesn't really match how humans sound.

You have to shape the voices in the tools, if you just spit them out they're junk but if you take the time to shape the voice a bit it gets better quickly, this is a cheap 11labs voice with 30 seconds spent on some basic shaping: https://s.h4x.club/bLuNlJWx Still a bit teleprompter-ish but there are tools to go in and adjust pace and style throughout and you probably hear a lot of stuff with people not using those cr…

What is shaping the voice means?

Re: PlayAI's new Dialog model achieves 3:1 preference in human evals

#26
post #25
post #21

Earlier quoted context omitted.

You have to shape the voices in the tools, if you just spit them out they're junk but if you take the time to shape the voice a bit it gets better quickly, this is a cheap 11labs voice with 30 seconds spent on some basic shaping: https://s.h4x.club/bLuNlJWx Still a bit teleprompter-ish but there are tools to go in and adjust pace and style throughout and you probably hear a lot of stuff with people not using those cr…

What is shaping the voice means?

You---can, really--slow, speed up or change, how, things sound, by, -- using queues like this, to control how the voice,,, - tells the story {{3sec}} - once you find a voice you like, you can go in and {{1sec}}

//

control how,

it goes about {{1sec}} story telling.

Re: PlayAI's new Dialog model achieves 3:1 preference in human evals

#27

For some reason, most of these (and other narration AIs) sound like someone reading off a teleprompter, rather than natural speaking voices. I'm not sure what exactly it is, but I'm left feeling like the speaker isn't really sure of what the next words are, and the stresses between the words are all over the place. It's like the emphasis over a sentence doesn't really match how humans sound.

Most of these models are trained on audiobooks, which could explain the teleprompter feeling vs a natural conversational feeling.

Re: PlayAI's new Dialog model achieves 3:1 preference in human evals

#29
post #15

I tried it with a paragraph of English taken from a formal speech, and it sounded quite good. I would not have been able to distinguish it from a skilled human narrator. But then I tried a paragraph of Japanese text, also from a formal speech, with the language set to Japanese and the narrator set to Yumiko Narrative. The result was a weird mixture of Korean, Chinese, and Japanese readings for the kanji and kana, all…

So kind of unrelated, but the reading/singing of arbitrary custom lyrics on suno.com's v4 model has blown me away.

suno is uncomfortably good. I run a group for helping founders and sometimes I make little suno songs to accompany the classes for fun, always impressed by what it spits out. (prompt: song for founder who have happy ears bringing them tears > 30 seconds gen >) https://s.h4x.club/p9u4ezl2 / https://s.h4x.club/mXuND7Eb / https://s.h4x.club/L1u2DYzW
Post reply on HN