Live data from Hacker News

Bagel: Open-source unified multimodal model

bagel-ai.org

31–35 of 35 posts

Re: Bagel: Open-source unified multimodal model

#31
post #19

A quick test in the "demo" link doesn't show it to be "as smart" as it appeared in the demos on the page. I really hope it does all it's promising to do, but I'm skeptic so far.

I found it surprising that even one of the demos on the page appeared to get it wrong. (Chat example #5, explaining the "My Handwriting In Exams" meme.) Not horribly wrong, but still an odd example to cherry-pick for publicity material.

ETA: oof, and it's still getting hands wrong. (Editing demo #12)

Re: Bagel: Open-source unified multimodal model

#32
post #15

I'm interested in potential alternatives to ChatGPT's advanced voice mode. When I see the word "multimodal" I'm hopeful the model understands text + voice but instead it almost always seems to refer to text + images. Is there a keyword that I can use to look for models that work with voice similar to ChatGPT's advanced voice mode?

I don't know that ChatGPT's voice mode is using audio as a transformer input directly.

It could just be using speech to text (e.g. Whisper) on your input, and then using its text model on the text of your words. Or has OpenAI said that they aren't doing this?

Re: Bagel: Open-source unified multimodal model

#34
post #15

I'm interested in potential alternatives to ChatGPT's advanced voice mode. When I see the word "multimodal" I'm hopeful the model understands text + voice but instead it almost always seems to refer to text + images. Is there a keyword that I can use to look for models that work with voice similar to ChatGPT's advanced voice mode?

I don't know that ChatGPT's voice mode is using audio as a transformer input directly. It could just be using speech to text (e.g. Whisper) on your input, and then using its text model on the text of your words. Or has OpenAI said that they aren't doing this?

OpenAI does not provide many details about their models these days but they do mention that the "Advanced voice" within ChatGPT operates on audio input directly:

> Advanced voice uses natively multimodal models, such as GPT-4o, which means that it directly “hears” and generates audio, providing for more natural, real-time conversations that pick up on non-verbal cues, such as the speed you’re talking, and can respond with emotion.

From https://help.openai.com/en/articles/8400625-voice-mode-faq

Re: Bagel: Open-source unified multimodal model

#35
post #20
post #15

I'm interested in potential alternatives to ChatGPT's advanced voice mode. When I see the word "multimodal" I'm hopeful the model understands text + voice but instead it almost always seems to refer to text + images. Is there a keyword that I can use to look for models that work with voice similar to ChatGPT's advanced voice mode?

Google Gemini Live is pretty good. If you want to try only voice, Try unmute.sh by Kyutai which will be eventually open-sourced

Thanks - it seems that Gemini Live is pretty far behind advanced voice mode at the moment. For example, I can't get it to speak slower when I want to understand what it is saying.

I'm still interested in what keyword I could use to search for the latest research in voice models.

Post reply on HN