Live data from Hacker News

VibeVoice: A Frontier Open-Source Text-to-Speech Model

microsoft.github.io

121–130 of 177 posts

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#121
post #90

Earlier quoted context omitted.

Different use cases: If you need a not-visual output of text, SoyA is a waste of electrons. If you want to try and mimic a human speaker, then it ain’t. Question is why would you need to have the computer sound more human, except for “because I can”.

I tried listening to audiobooks generated with tts. It takes me out of it most of the time, and I lose focus. That podcast thing from google was the first time I felt like I could listen to an entire thing without feeling the uncanny valley thing. And I knew it was genAI. So I'm looking for that, but for my content. Grab a bunch of articles (long form, deeply researched) and "podcast" them but with natural voices, sa…

The Google podcasts are so cringey positive it emotionally pains me. Nobody finds pineapple on pizza that amazing.

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#122

What an odd name to me, becaus "Vibe" is, in my mind, equal to somewhat poor quality. Like "Vibe Coding". But that's probably just some bias from my side.

Vibe always meant "specific feel" and makes sense related to AI coding "by touch" vs. understanding what's actually happening. It's just the results have now made the word pejorative.

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#124
post #76

Earlier quoted context omitted.

Is there any better model you can point at? I would be interested in having a listen. There are people – and it does not matter what it's about – that will overstate the progress made (and others will understate it, case in point). Neither should put a damper on progress. This is the best I personally have heard so far, but I certainly might have missed something.

It’s tough to name the best local TTS since they all seem to trade off on quality and features and none of them are as good as ElevenLabs’ closed-source offering. However Kokoro-82M is an absolute triumph in the small model space. It curbstomps models 10-20x its size in terms of quality while also being runnable on like, a Raspberry Pi. It’s the kind of thing I’m surprised even exists. Its downside is that it isn’t s…

What is your opinion about F5-TTS or Fish-TTS?

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#125

I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…

The male Chinese speakers had THICK American accents. Nothing really wrong with the language, but think the stereotype German speaking English. That was kind of strange to me.

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#126

I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…

Is there any better model you can point at? I would be interested in having a listen. There are people – and it does not matter what it's about – that will overstate the progress made (and others will understate it, case in point). Neither should put a damper on progress. This is the best I personally have heard so far, but I certainly might have missed something.

Probably not even the best ones, but among some recent models I find Dia and Orpheus more natural

- http://dia-tts.com/

- https://github.com/canopyai/Orpheus-TTS

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#127
post #125

I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…

The male Chinese speakers had THICK American accents. Nothing really wrong with the language, but think the stereotype German speaking English. That was kind of strange to me.

I think it's because it was using the American voice for it. Conversely the female voice in the Mandarin conversation spoke English with a Chinese accent.

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#128

Earlier quoted context omitted.

I tried listening to audiobooks generated with tts. It takes me out of it most of the time, and I lose focus. That podcast thing from google was the first time I felt like I could listen to an entire thing without feeling the uncanny valley thing. And I knew it was genAI. So I'm looking for that, but for my content. Grab a bunch of articles (long form, deeply researched) and "podcast" them but with natural voices, sa…

The Google podcasts are so cringey positive it emotionally pains me. Nobody finds pineapple on pizza that amazing.

>Nobody finds pineapple on pizza that amazing

We can't be friends

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#129
post #75

Earlier quoted context omitted.

The most popular song of the year from one of the most popular movie franchises that had been in the global news due to the death of its star. Probably the most memorable song from a soundtrack of the century so far.

I'm Just Ken (Barbie), Skyfall, Let it Go (Frozen), Remember Me (Coco), Happy (from Despicable Me 2), a Star is Born (Shallow), are all arguably wayyyyy more memorable and these are just off the top of my head. We've had quite a few memorable songs in soundtracks this millennium. edit: I had forgotten about Jai Ho (Slumdog Millionaire) and Lose Yourself (8 mile)

And most recently "Golden"

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#130

I really hope someone within Microsoft is naming their open source coding agent Microsoft VibeCode. Let this be a thing. Its either that or "Lo" then you can have Lo work with Phi, so you can Vibe code with Lo Phi. https://techcommunity.microsoft.com/blog/azure-ai-foundry-bl...

[deleted]
Post reply on HN