Open-source, eh? Where's the training data, then?
Most scraped data is often full of copyright, usage agreement, and privacy law violations. Making it "open" would be unwise for a commercial entity. =3
VibeVoice: A Frontier Open-Source Text-to-Speech Model
141–150 of 177 posts
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#142Earlier quoted context omitted.
Most scraped data is often full of copyright, usage agreement, and privacy law violations. Making it "open" would be unwise for a commercial entity. =3
Open source is being abused to not provide the actual source. Stop this.
For example, many academic data sets are not public domain, and can't be used in a commercial context. A GPL claim on that data is often an argument of which thief showed up first.
Rule #24: A lawyers Strategic Truth is to never lie, but also avoid voluntarily disclosing information that may help opponents.
Thus, a business will never disclose they paid a fool to break laws for them... =3
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#143I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…
Is there any better model you can point at? I would be interested in having a listen. There are people – and it does not matter what it's about – that will overstate the progress made (and others will understate it, case in point). Neither should put a damper on progress. This is the best I personally have heard so far, but I certainly might have missed something.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#144Earlier quoted context omitted.
Their comments about the singing and background music are odd. It’s been a while since I’ve done academic research, but something about those comments gave me a strong “we couldn’t figure out how to make background music go away in time for our paper submission, so we’re calling it a feature” vibe as opposed to a “we genuinely like this and think its a differentiator” vibe.
Totally felt the same way! Singing happens spontaneously? What?
> In fact, we intentionally decided not to denoise our training data because we think it's an interesting feature for BGM to show up at just the right moment. You can think of it as a little easter egg we left for you.
It's not a bug, it's a feature! Okaaaaay
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#145Earlier quoted context omitted.
There's a lot of money and effort spent in satisfying the sexual desires of (predominantly straight) men. There's not typically quite as much interest in doing the same for women. For example I've been looking at models and loras for generating images, and the boards are _full_ of ones that will generate women well or in some particular style. Quite often at least a couple of the preview images for each are hidden be…
I think this is a very lazy kind of cultural analysis. The reason female voices are being chosen over male ones is a little more multifaceted than just SEX. Heterosexual women also tend to prefer female voices over male ones. Female voices are often rated as being clearer, easier to understand, "warmer", etc. Why this is the case is still an open question, but it's definitely more complex than just SEX.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#146I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…
It's good but not the best free model. I find Chatterbox to be more realistic with no robot-sounding and better (though not perfect) intonation.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#147Earlier quoted context omitted.
I think this is a very lazy kind of cultural analysis. The reason female voices are being chosen over male ones is a little more multifaceted than just SEX. Heterosexual women also tend to prefer female voices over male ones. Female voices are often rated as being clearer, easier to understand, "warmer", etc. Why this is the case is still an open question, but it's definitely more complex than just SEX.
That you consider it sex (rather than gender), is exactly why there’s a preference for female coded voices. Consider where we do hear male recorded voices used as default.
> satisfying the sexual desires of
So, "sex" as a reference to "sexual desires". In English, it just so happens that "sex" has other meanings, but those weren't in play at the time.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#148Earlier quoted context omitted.
Open source is being abused to not provide the actual source. Stop this.
A lot of code has multiple FOSS licenses that are not contaminating like GPL. GPL violations do occur on code, but have nothing to do with the training Data. For example, many academic data sets are not public domain, and can't be used in a commercial context. A GPL claim on that data is often an argument of which thief showed up first. Rule #24: A lawyers Strategic Truth is to never lie, but also avoid voluntarily d…
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#149Earlier quoted context omitted.
Open source is being abused to not provide the actual source. Stop this.
A lot of code has multiple FOSS licenses that are not contaminating like GPL. GPL violations do occur on code, but have nothing to do with the training Data. For example, many academic data sets are not public domain, and can't be used in a commercial context. A GPL claim on that data is often an argument of which thief showed up first. Rule #24: A lawyers Strategic Truth is to never lie, but also avoid voluntarily d…
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#150Earlier quoted context omitted.
A lot of code has multiple FOSS licenses that are not contaminating like GPL. GPL violations do occur on code, but have nothing to do with the training Data. For example, many academic data sets are not public domain, and can't be used in a commercial context. A GPL claim on that data is often an argument of which thief showed up first. Rule #24: A lawyers Strategic Truth is to never lie, but also avoid voluntarily d…
Perhaps, but it is not Open Source in the traditional sense if they do not provide the preferred form for modifications.
Indeed, these adversarial behaviors do not follow the spirit of FOSS community standards. If a project started as FOSS, than FOSS it should remain. =3