Live data from Hacker News

VibeVoice: Open-source frontier voice AI

github.com

41–50 of 191 posts

Re: VibeVoice: Open-source frontier voice AI

#41

Earlier quoted context omitted.

> we should stop calling this type of model open source. They are indeed "open weight” This ship has sailed. It’s now in the same category as hacker/cracker and the pronunciation of GIF.

It's the same as GIS, you wouldn't say jizz now would you?

I absolutely do, every single time it comes up.

Re: VibeVoice: Open-source frontier voice AI

#44
post #14

I think we should stop calling this type of models open source. They are indeed "open weight." The training code is proprietary and never revealed. https://github.com/microsoft/VibeVoice/issues/102

Openwashing is the new greenwashing, which, coincidently, seems to have gone out of fashion a few hundred datacentres ago.

it was replaced with abundancewashing

Re: VibeVoice: Open-source frontier voice AI

#45

Earlier quoted context omitted.

> we should stop calling this type of model open source. They are indeed "open weight” This ship has sailed. It’s now in the same category as hacker/cracker and the pronunciation of GIF.

It's the same as GIS, you wouldn't say jizz now would you?

The developer of the format declared the pronunciation 30+ years ago. It has always been jif.

Re: VibeVoice: Open-source frontier voice AI

#46
I took a look into local options for ASR and diarization some months ago, I missed that VibeVoice now has this feature.

My conclusions back then (which only came from a shallow research on the topic and 0 real experience mind you) was that Whisper + Pyannote was the "stable" approach.

Have the VibeVoice, Voxtral, Qwen or the Nemo solutions caught up in segmentation and speaker recognition?

Re: VibeVoice: Open-source frontier voice AI

#47
post #13

I the past month or so, I added 2 models to my app Whisper Memos ( https://whispermemos.com ): - Cohere Transcribe (self hosted) - Grok Speech To Text (they provide an API, only $0.10/hr!) They are both excellent. I'm not sure about this one. Would you like to see it in a consumer speech to text app?

Does Cohere work with longer transcripts? Do you have to do some magic to merge recordings over 35 seconds long?

Re: VibeVoice: Open-source frontier voice AI

#48
post #14

I think we should stop calling this type of models open source. They are indeed "open weight." The training code is proprietary and never revealed. https://github.com/microsoft/VibeVoice/issues/102

Indeed. We now live in a world where freeware is named open source. We are very sorry, Stallman.

Re: VibeVoice: Open-source frontier voice AI

#49
post #34
post #28

Interesting to see "vibe" enshrined by the likes of Microsoft as an AI product word.

Especially when "vibe coded" can have a negative connotation meaning quickly put together without understanding.

I’m just surprised they put the name of the e-waste slop company in their product

Re: VibeVoice: Open-source frontier voice AI

#50
post #13

I the past month or so, I added 2 models to my app Whisper Memos ( https://whispermemos.com ): - Cohere Transcribe (self hosted) - Grok Speech To Text (they provide an API, only $0.10/hr!) They are both excellent. I'm not sure about this one. Would you like to see it in a consumer speech to text app?

Any non-Musk alternatives that are comparable in quality and cost?

Voxtral competes on price ($0.003/min) and quality. Speechmatics has best in class accuracy but is a bit more expensive ($0.004/min)
Post reply on HN