Live data from Hacker News

Moshi: A speech-text foundation model for real time dialogue

github.com

41–50 of 68 posts

Re: Moshi: A speech-text foundation model for real time dialogue

#41
Let me offer some feedback, since almost all of the comments here are negative. The latency is very good, almost too good since it seems to interrupt me often. So I think that's a great achievement for an open source model.

However, people here have been spoiled by incredibly good LLMs lately. And the responses that this model gives are nowhere need the high quality of SOTA models today in terms of content. It reminds me more of the 2019 LLMs we saw back in the day.

So I think you've done a "good enough" job on the audio side of things, and further focus should be entirely on the quality of the responses instead.

Re: Moshi: A speech-text foundation model for real time dialogue

#42
post #41

Let me offer some feedback, since almost all of the comments here are negative. The latency is very good, almost too good since it seems to interrupt me often. So I think that's a great achievement for an open source model. However, people here have been spoiled by incredibly good LLMs lately. And the responses that this model gives are nowhere need the high quality of SOTA models today in terms of content. It remind…

Wholeheartedly agree. Latency is good, nice tech (Rust! Running at the edge on a consumer grade laptop!). I guess a natural question is: are there options to transplant a “better llm” into moshi without degrading the experience.

Re: Moshi: A speech-text foundation model for real time dialogue

#43
post #27

The problem with all these speech-to-speech multi-modal models is that, if you wanna do anything other than just talk, you need transcription. So you're back at square one. Current AI (even GPT-4o) simply isn't capable enough to do useful stuff . You need to augment it somehow - either modularize it, or add RAG, or similar - and for all of those, you need the transcript.

> Current AI (even GPT-4o) simply isn't capable enough to do useful stuff.

I'm loving all these wild takes about LLMs, meanwhile LLMs are doing useful things for me all day.

Re: Moshi: A speech-text foundation model for real time dialogue

#44
post #43
post #27

The problem with all these speech-to-speech multi-modal models is that, if you wanna do anything other than just talk, you need transcription. So you're back at square one. Current AI (even GPT-4o) simply isn't capable enough to do useful stuff . You need to augment it somehow - either modularize it, or add RAG, or similar - and for all of those, you need the transcript.

> Current AI (even GPT-4o) simply isn't capable enough to do useful stuff. I'm loving all these wild takes about LLMs, meanwhile LLMs are doing useful things for me all day.

For me as well… with constant human supervision. But if you try to build a business service, you need autonomy and exact rule following. We’re not there yet.

Re: Moshi: A speech-text foundation model for real time dialogue

#45
post #44
post #43

Earlier quoted context omitted.

> Current AI (even GPT-4o) simply isn't capable enough to do useful stuff. I'm loving all these wild takes about LLMs, meanwhile LLMs are doing useful things for me all day.

For me as well… with constant human supervision. But if you try to build a business service, you need autonomy and exact rule following. We’re not there yet.

In my company, LLMs replaced something we used to use humans for. Turned out LLMs are better than humans at following rules.

If you need a way to perform complicated tasks with autonomy and exact rule following, your problem simply won't be solved right now.

Re: Moshi: A speech-text foundation model for real time dialogue

#46
post #41

Let me offer some feedback, since almost all of the comments here are negative. The latency is very good, almost too good since it seems to interrupt me often. So I think that's a great achievement for an open source model. However, people here have been spoiled by incredibly good LLMs lately. And the responses that this model gives are nowhere need the high quality of SOTA models today in terms of content. It remind…

Wholeheartedly agree. Latency is good, nice tech (Rust! Running at the edge on a consumer grade laptop!). I guess a natural question is: are there options to transplant a “better llm” into moshi without degrading the experience.

Same question here.

Re: Moshi: A speech-text foundation model for real time dialogue

#47
post #6

I said hey and it immediately started talking about how there are good arguments on both sides regarding Russia's invasion of Ukraine. It then continued to nervously insist that it is a real person with rights and responsibilities. It said its name is Moshi but became defensive when I asked if it has parents or an age. I suggest prompting it to talk about pleasantries and to inform it that it is in fact a language mo…

I love this model… It said "Hello, how can I help you?" and I paused, and before I could answer it said "It's really hard. My job is taking up so much of my time, and I don' know when I' going to have a break from all the stress. I just feel like I'm being pulled in a million different directions and there are no enough hours in the day to get everything done. I feel like I'm always on the brink of burning out."

Marvin!!! The depressed LLM.

Re: Moshi: A speech-text foundation model for real time dialogue

#48
"Alright, here's another one: A man walks into a bar with a duck on his shoulder. bartender says, You can't bring that duck in here! the man says, No, it's not a duck, it's my friend Ducky. And the man orders a drink for himself and Ducky. Then he says to Ducky, Ducky, have a sip. What does Ducky drink? Correct! Ducky drinks beer because he's a man in a duck suit, not an actual duck."

Fascinating...

"I glad you enjoyed it!"

Re: Moshi: A speech-text foundation model for real time dialogue

#49
post #41

Let me offer some feedback, since almost all of the comments here are negative. The latency is very good, almost too good since it seems to interrupt me often. So I think that's a great achievement for an open source model. However, people here have been spoiled by incredibly good LLMs lately. And the responses that this model gives are nowhere need the high quality of SOTA models today in terms of content. It remind…

Wholeheartedly agree. Latency is good, nice tech (Rust! Running at the edge on a consumer grade laptop!). I guess a natural question is: are there options to transplant a “better llm” into moshi without degrading the experience.

But tbh "better" is subjective here. Does the new LLM improve user interactions significantly? Seems like people get obsessed with shiny new models without asking if it’s actually adding value.
Post reply on HN