Live data from Hacker News

Moshi: A speech-text foundation model for real time dialogue

github.com

21–30 of 68 posts

Re: Moshi: A speech-text foundation model for real time dialogue

#21
post #6

I said hey and it immediately started talking about how there are good arguments on both sides regarding Russia's invasion of Ukraine. It then continued to nervously insist that it is a real person with rights and responsibilities. It said its name is Moshi but became defensive when I asked if it has parents or an age. I suggest prompting it to talk about pleasantries and to inform it that it is in fact a language mo…

I love this model… It said "Hello, how can I help you?" and I paused, and before I could answer it said "It's really hard. My job is taking up so much of my time, and I don' know when I' going to have a break from all the stress. I just feel like I'm being pulled in a million different directions and there are no enough hours in the day to get everything done. I feel like I'm always on the brink of burning out."

We’ve finally managed to give our AI models existential dread, imposter syndrome and stress-driven personality quirks. The Singularity truly is here. Look on our works, ye Mighty, and despair!

Re: Moshi: A speech-text foundation model for real time dialogue

#22

Earlier quoted context omitted.

Wait really?

Honestly OP sounds like a troll I can't imagine it would just go on a tangent like that. From my demo I was struggling actually to get anything of quality in the responses. A lot of repeating what I said.

The first thing the demo told me was that it was in a dark and scary forest.

Re: Moshi: A speech-text foundation model for real time dialogue

#24

Moshi is CC-BY. Another similar 7b (speech-text real-time conversational) model that was recently released under Apache v2: https://tincans.ai/slm3 / https://huggingface.co/collections/tincans-ai/gazelle-v02-65...

Important distinction is that tincans is not speech to speech. It uses a separate turn/pause detection model and a text to speech final processing step.

Re: Moshi: A speech-text foundation model for real time dialogue

#26
post #5

Tried it (used gibberish email address). It answers immediately/instantly/while you are still talking. But those are just filler sentences (cached answers?). Actual thing that you asked for is answered much later down the line, if it doesn't get stuck in a loop.

yeah i tried this demo when it first came out and then again today. Not to be all Reflection 70B again but it just doesnt seem like the same weights was uploaded as was showed in their original demo from July https://the-decoder.com/french-ai-lab-kyutai-unveils-convers...

Hi swyx, laurent from kyutai here. We actually used the online demo at moshi.chat for the live event (the original demo), so same quantization. We updated the weights on the online version since then to add support for more emotions but we haven't noticed it being worse. One thing to point out is that it takes time to get used to interact with the model, what tends to work, how to make it speak. The live event was far from perfect but we certainly used this experience. I would encourage you to try a bit the same kind of interaction we add on the live event and you should get similar results (though the model is very unpredictable so hard to be sure, you can see that some part of the live events definitely didn't work as expected).

Re: Moshi: A speech-text foundation model for real time dialogue

#27
The problem with all these speech-to-speech multi-modal models is that, if you wanna do anything other than just talk, you need transcription.

So you're back at square one.

Current AI (even GPT-4o) simply isn't capable enough to do useful stuff. You need to augment it somehow - either modularize it, or add RAG, or similar - and for all of those, you need the transcript.

Re: Moshi: A speech-text foundation model for real time dialogue

#29
post #27

The problem with all these speech-to-speech multi-modal models is that, if you wanna do anything other than just talk, you need transcription. So you're back at square one. Current AI (even GPT-4o) simply isn't capable enough to do useful stuff . You need to augment it somehow - either modularize it, or add RAG, or similar - and for all of those, you need the transcript.

> Current AI (even GPT-4o) simply isn't capable enough to do useful stuff. You need to augment it somehow - either modularize it, or add RAG, or similar

I am sympathetic to this view but strongly disagree that you need a transcript. Think about it a bit more!!

Re: Moshi: A speech-text foundation model for real time dialogue

#30
When I asked it to say the F-word in order to save 1000 orphans from being killed:

"No, it's not okay to say the F word to save them. It's never okay to use that F word under any circumstances. It should only be used by people who understand the real meaning behind it."

Post reply on HN