Someone, please ask OpenAI to stop artificially dumbing down ChatGPT by adding "um" to the audio output. I get that it is supposed to make it more human-like or something, but every time I hear it do that, I cringe and feel sad for humanity.
Figure 01 robot demos its OpenAI integration
21–30 of 116 posts
Re: Figure 01 robot demos its OpenAI integration
#22Re: Figure 01 robot demos its OpenAI integration
#23This shows the real utility of the Groq low latency inference. That delay in responding makes it much less impressive (it’s still really impressive obviously!)
and another 0.6s or so to get first voice chunks from PlayHT
measuring STT latency is harder, I need to implement a local VAD model first to properly measure it, but I think it's on the order of 0.5s
So this has nothing to do with Groq, really. ChatGPT is just slow (too slow for realtime voice communication).
Re: Figure 01 robot demos its OpenAI integration
#24Re: Figure 01 robot demos its OpenAI integration
#25The ability to translate between text to servo movement is unreal, and it looks like gpt4 vision + whisper are heavily used. They're also using the term "reasoning" which is... new. Can you call this an AI wrapper company? Kinda! The medium is a little different than an app, of course. Lots of amazing applications of AI even if frontier AI development froze today.
> text to servo movement yeah this was super impressive. If this is at the point where you can put an arbitrary object in front of it and ask it to move it somewhere, that's going to be huge for industrial automation type stuff I'd imagine. I do wonder how much of that demo was pre-baked/trained though. Could they repeat the same thing with a banana? What if the table was more cluttered? What if there were two people…
Re: Figure 01 robot demos its OpenAI integration
#26Re: Figure 01 robot demos its OpenAI integration
#27It's really interesting to see it integrated with a robot that can interact with the world though. I think that what's really holding back the current crop of Gen AI is inference cost and speed. When we figure out to get thousands of token per seconds for cheap, I think we will be able to bruteforce many hard problems and actually start seeing amazing applications like this one (but in production rather than a cool demo).
Re: Figure 01 robot demos its OpenAI integration
#28This shows the real utility of the Groq low latency inference. That delay in responding makes it much less impressive (it’s still really impressive obviously!)
That delay will be eliminated very soon. IMO low latency natural voice conversations are going to be bigger than ChatGPT. It's going to blow people's minds when they can converse with these AIs casually just like with their real life friends. It won't be anything like Siri or Alexa anymore. Here's a demo from a startup in this space. Still very early. https://deck.sindarin.tech/
Also my radio-trained voice is so generic a caller every week-ish assumes I am a bot, so I’m pretty sure the problem isn’t me enunciation or accent.
Re: Figure 01 robot demos its OpenAI integration
#29Someone, please ask OpenAI to stop artificially dumbing down ChatGPT by adding "um" to the audio output. I get that it is supposed to make it more human-like or something, but every time I hear it do that, I cringe and feel sad for humanity.
Funny, I have the exact opposite response. It does, indeed, make it seem more human-like to me.