Live data from Hacker News

PlayAI's new Dialog model achieves 3:1 preference in human evals

play.ht

31–40 of 59 posts

Re: PlayAI's new Dialog model achieves 3:1 preference in human evals

#31

For some reason, most of these (and other narration AIs) sound like someone reading off a teleprompter, rather than natural speaking voices. I'm not sure what exactly it is, but I'm left feeling like the speaker isn't really sure of what the next words are, and the stresses between the words are all over the place. It's like the emphasis over a sentence doesn't really match how humans sound.

Yup, and that's going to be the case until AI's can really model human psychology. Speech encodes a gigantic amount of emotion via prosody and rhythm -- how the speaker is feeling, how they feel about each noun and verb, what they're trying to communicate with it. If you try to reproduce all the normal speech prosody, it'll be all over the place and SoUnD bIzArRe and won't make any sense, and be incredibly distractin…

What if you have it read the script, then say, “hey, at this point, what is the character feeling? What are they trying to accomplish? What is there relationship to each person in the scene?”

And then you get that and prompt the model to add inflection and pacing and whatever to the text to reflect that. You feed that into the speech model.

It seems like it could definitely do the first part (“based on this text, this character might be feeling X”); the second part (“mark up the dialogue”) seems easier; the third part about speech seems doable already based on another comment.

So we are pretty close already? Whatever actors are doing can be approximated through prompting, including the director iterating with the “actors”.

Re: PlayAI's new Dialog model achieves 3:1 preference in human evals

#35
post #28

I love the tech. I hate that it gets used to fill YouTube with zero-effort slop. I don't have a solution.

I realize it’s hard to face, but it’s ok to admit that (cool as the tech is) some things just have a net negative on the world. It’s just engineers and data scientists using their enormous talents to make world a worse place, instead of a better one.

It’s not going to happen, but the only solution is to just stop developing it.

Re: PlayAI's new Dialog model achieves 3:1 preference in human evals

#36
post #35
post #28

I love the tech. I hate that it gets used to fill YouTube with zero-effort slop. I don't have a solution.

I realize it’s hard to face, but it’s ok to admit that (cool as the tech is) some things just have a net negative on the world. It’s just engineers and data scientists using their enormous talents to make world a worse place, instead of a better one. It’s not going to happen, but the only solution is to just stop developing it.

[dead]

Re: PlayAI's new Dialog model achieves 3:1 preference in human evals

#37
post #35
post #28

I love the tech. I hate that it gets used to fill YouTube with zero-effort slop. I don't have a solution.

I realize it’s hard to face, but it’s ok to admit that (cool as the tech is) some things just have a net negative on the world. It’s just engineers and data scientists using their enormous talents to make world a worse place, instead of a better one. It’s not going to happen, but the only solution is to just stop developing it.

The cat's out of the bag, if someone stops then someone else would start.

I would entreat people to consider the net effect of anything they create. Let it at least sway your decisions somewhat. It probably won't be enough to not do it, but I think of it more as the ratio between net positive :: net negative, and paying attention to that ratio should help swing it at least somewhat -- certainly more than giving up and ignoring the benefits :: harms.

Re: PlayAI's new Dialog model achieves 3:1 preference in human evals

#38
post #35
post #28

I love the tech. I hate that it gets used to fill YouTube with zero-effort slop. I don't have a solution.

I realize it’s hard to face, but it’s ok to admit that (cool as the tech is) some things just have a net negative on the world. It’s just engineers and data scientists using their enormous talents to make world a worse place, instead of a better one. It’s not going to happen, but the only solution is to just stop developing it.

one issue is that having a good generative model is unavoidable as a component for lots of good, useful tasks. like translation, transcription, etc.

but then of course if you have a generative model, you can use it to generate stuff.

Re: PlayAI's new Dialog model achieves 3:1 preference in human evals

#39
To the founders: Would love to share my audio files with my team before we commit to a payment plan. Is there anyway to share audio files I've generated?

Our whole team is on Elevenlabs and a switch is significant work, but I think the results are worth it! Super awesome work!

Re: PlayAI's new Dialog model achieves 3:1 preference in human evals

#40
post #29

Earlier quoted context omitted.

So kind of unrelated, but the reading/singing of arbitrary custom lyrics on suno.com's v4 model has blown me away.

suno is uncomfortably good. I run a group for helping founders and sometimes I make little suno songs to accompany the classes for fun, always impressed by what it spits out. (prompt: song for founder who have happy ears bringing them tears > 30 seconds gen >) https://s.h4x.club/p9u4ezl2 / https://s.h4x.club/mXuND7Eb / https://s.h4x.club/L1u2DYzW

Suno songs always have way too much treble or reverb, or something I can't quite put my finger on. They're very bright sounding.

I don't think it's a fatal flaw, but I hope future versions improve on this, or Suno starts doing some more post-processing to address it. I know there's a new "remaster" feature, but I'm not sure if it does anything there either.

Post reply on HN