Live data from Hacker News

VibeVoice: A Frontier Open-Source Text-to-Speech Model

microsoft.github.io

91–100 of 177 posts

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#92
post #32

I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…

One of the things this model is actually quite good at is voice cloning. Drop a recorded sample of your voice into the voices folder, and it just works.

bonus usage

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#93

I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…

ElevenLabs has a much more convincing voice model

it's not oss

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#94

I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…

Is there any better model you can point at? I would be interested in having a listen. There are people – and it does not matter what it's about – that will overstate the progress made (and others will understate it, case in point). Neither should put a damper on progress. This is the best I personally have heard so far, but I certainly might have missed something.

I cobbled together llm-tts to run as many local (and remote) TTs models s I could find and get working.

https://github.com/mlang/llm-tts

Strictly speaking, even music generation fits the usage pattern: text in, audio out.

llm-tts is far from complete, but it makes it relatively "easy" to try a few models in an uniform way.

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#96
post #90

Earlier quoted context omitted.

> with acceptable quality Compared to IBMs Steven Hawking's chair, maybe. But apple tts is not acceptable quality in any modern understanding of SotA, IMO.

Different use cases: If you need a not-visual output of text, SoyA is a waste of electrons. If you want to try and mimic a human speaker, then it ain’t. Question is why would you need to have the computer sound more human, except for “because I can”.

I tried listening to audiobooks generated with tts. It takes me out of it most of the time, and I lose focus. That podcast thing from google was the first time I felt like I could listen to an entire thing without feeling the uncanny valley thing. And I knew it was genAI. So I'm looking for that, but for my content. Grab a bunch of articles (long form, deeply researched) and "podcast" them but with natural voices, sans hype. Or books. Have them ready when I'm out and about.
Post reply on HN