Live data from Hacker News

PlayAI's new Dialog model achieves 3:1 preference in human evals

play.ht

41–50 of 59 posts

Re: PlayAI's new Dialog model achieves 3:1 preference in human evals

#41
I wanted to use this to do temp voices for a video game project. Not realtime, just creating the audio at build time basically. However, the pricing model is not conducive to that because you cannot pay-per-use, and on top of that none of the lower cost plans support more than 10 requests per minute so its difficult to use for batch operations. $299/mo seemed steep for my use case of infrequent bursts, and they couldn't help me with a custom plan, so I have ended up just using Azure AI Text-to-Speech. (Which is also much faster to render.)

Re: PlayAI's new Dialog model achieves 3:1 preference in human evals

#42
post #29

Earlier quoted context omitted.

suno is uncomfortably good. I run a group for helping founders and sometimes I make little suno songs to accompany the classes for fun, always impressed by what it spits out. (prompt: song for founder who have happy ears bringing them tears > 30 seconds gen >) https://s.h4x.club/p9u4ezl2 / https://s.h4x.club/mXuND7Eb / https://s.h4x.club/L1u2DYzW

Suno songs always have way too much treble or reverb, or something I can't quite put my finger on. They're very bright sounding. I don't think it's a fatal flaw, but I hope future versions improve on this, or Suno starts doing some more post-processing to address it. I know there's a new "remaster" feature, but I'm not sure if it does anything there either.

Yeah they're way too wide and not muddy enough, if you're gonna be as wide as they often are you need to fill it, else they always just sound over produced/weirdly bright. I was thinking earlier about why they sound really good but not actually good and then I realized that's how I feel about most modern pop music anyway. I think the main thing you're hearing, or at least the thing I find annoying, is if you listen close to how the AI does harmony it seems to almost be cloning the original vocal line and pitching it up and out so it's slightly offset feeling giving the appearance of a second vocalist, its a tick I do in abelton to see what things might sound like build differently with vocals and it feels very much like the sumo fake backing singers. I do think they're like 6 months or so away from nailing a lot of this given how quickly they've been moving, I follow them closely and it's been impressive. (you can also pull meatier stuff out of it if you work it a bit: https://s.h4x.club/04uz6klg - don't think this sounds particularly "AI" at all - edit: turns out if you play up in chamber orchestra and choir in the prompting you can get some much better stuff out of it: https://s.h4x.club/eDubr9xJ)

Re: PlayAI's new Dialog model achieves 3:1 preference in human evals

#43
Maybe launched a bit too quickly? You can select "Swedish" for the non-Swedish voices, but the results are very poor. Far from useable. And there is no Swedish voice. So that language support claim is made a bit to soon I would say.

Also I found no way to filter/sort the voice selection modal on language, so I have to visually search the entire list.

Re: PlayAI's new Dialog model achieves 3:1 preference in human evals

#44
post #15

I tried it with a paragraph of English taken from a formal speech, and it sounded quite good. I would not have been able to distinguish it from a skilled human narrator. But then I tried a paragraph of Japanese text, also from a formal speech, with the language set to Japanese and the narrator set to Yumiko Narrative. The result was a weird mixture of Korean, Chinese, and Japanese readings for the kanji and kana, all…

Alternatively, text that is input to these services should be passed through a normalization process, i.e. use LLAMA to convert kanji to hiragana or a romanization. The TTS output is then much better.

Re: PlayAI's new Dialog model achieves 3:1 preference in human evals

#45
post #37
post #35

Earlier quoted context omitted.

I realize it’s hard to face, but it’s ok to admit that (cool as the tech is) some things just have a net negative on the world. It’s just engineers and data scientists using their enormous talents to make world a worse place, instead of a better one. It’s not going to happen, but the only solution is to just stop developing it.

The cat's out of the bag, if someone stops then someone else would start. I would entreat people to consider the net effect of anything they create. Let it at least sway your decisions somewhat. It probably won't be enough to not do it, but I think of it more as the ratio between net positive :: net negative, and paying attention to that ratio should help swing it at least somewhat -- certainly more than giving up an…

The whole idea of developing AGI (even if LLMs are probably the wrong approach) is so strange when you think about it.

The smartest people in the world are working very hard in order to make themselves completely redundant and cheaply replaceable. If they succeed, they will turn their main skill and defining characteristic into a meaningless curiosity.

And life will not be better, the manual work of today will still need to be done and robots are not up to the task. Even with better programming for cheap, the hardware (and I mean even simply the metal) is too expensive compared to human hands.

Re: PlayAI's new Dialog model achieves 3:1 preference in human evals

#46
post #15

I tried it with a paragraph of English taken from a formal speech, and it sounded quite good. I would not have been able to distinguish it from a skilled human narrator. But then I tried a paragraph of Japanese text, also from a formal speech, with the language set to Japanese and the narrator set to Yumiko Narrative. The result was a weird mixture of Korean, Chinese, and Japanese readings for the kanji and kana, all…

Alternatively, text that is input to these services should be passed through a normalization process, i.e. use LLAMA to convert kanji to hiragana or a romanization. The TTS output is then much better.

Unfortunately, a simple normalization of kanji --> hiragana throws away pronunciation information.

Re: PlayAI's new Dialog model achieves 3:1 preference in human evals

#47
post #29

Earlier quoted context omitted.

So kind of unrelated, but the reading/singing of arbitrary custom lyrics on suno.com's v4 model has blown me away.

suno is uncomfortably good. I run a group for helping founders and sometimes I make little suno songs to accompany the classes for fun, always impressed by what it spits out. (prompt: song for founder who have happy ears bringing them tears > 30 seconds gen >) https://s.h4x.club/p9u4ezl2 / https://s.h4x.club/mXuND7Eb / https://s.h4x.club/L1u2DYzW

AI generated music is like AI Art.

It feels really generic. But to be fair a lot of art is just like that.

How many animated shows use the Family Guy laziest common denominator style. Storylines that are written to be easy to follow and mundane.

Ask Chat GPT to write a complex story about divorce and trama. It'll either refuse to it or come up with a Hallmark ending.

Re: PlayAI's new Dialog model achieves 3:1 preference in human evals

#48
post #20

I've been messing with the open source side of audio generation, and expressiveness still takes work but it's getting there. Roughly summarized my findings are: - zero shot voice cloning isn't there yet - gpt-sovits is the best at non-word vocalizations, but the overall quality is bad when just using zero shot, finetuning helps - F5 and fish-speech are both good as well - xtts for me has had the best stability (i can…

What scenario would be considered "zero shot" voice cloning? At the very least, wouldn't you have to provide 1 sample? Which would make it "few shot" (if that term really even makes sense in this context).

I think the key distinction is that there is no specific training data for that speaker. You can view the input as just the input voice to clone, not training examples.

It would be more like training examples if you had to give it specific phrases.

Re: PlayAI's new Dialog model achieves 3:1 preference in human evals

#49
post #29

Earlier quoted context omitted.

suno is uncomfortably good. I run a group for helping founders and sometimes I make little suno songs to accompany the classes for fun, always impressed by what it spits out. (prompt: song for founder who have happy ears bringing them tears > 30 seconds gen >) https://s.h4x.club/p9u4ezl2 / https://s.h4x.club/mXuND7Eb / https://s.h4x.club/L1u2DYzW

AI generated music is like AI Art. It feels really generic. But to be fair a lot of art is just like that. How many animated shows use the Family Guy laziest common denominator style. Storylines that are written to be easy to follow and mundane. Ask Chat GPT to write a complex story about divorce and trama. It'll either refuse to it or come up with a Hallmark ending.

I mean yes and no, like if you just let it generate lazily, then yes. However if you work on lyrics and generate a bunch of samples.. no it can be very powerful and artistic.
Post reply on HN