Voice has the potential to be awesome. This demo is really underwhelming to me because of the multi-second latency between the query and response, just like every other lame voice assistant. It doesn't have to be this way! I have a local demo using Llama 2 that responds in about half a second and it feels like talking to an actual person instead of like Siri or something. I really should package it up so people can t…
Completely agree, latency is key for unlocking great voice experiences. Here's a quick demo I'm working on for voice ordering https://youtu.be/WfvLIEHwiyo Total end-to-end latency is a few hundred milliseconds: starting from speech to text, to the LLM, then to a POS to validate the SKU (no hallucinations are possible!), and finally back to generated speech. The latency is starting to feel really natural. Building out…
We are beginning to roll out new voice and image capabilities in ChatGPT
641–650 of 914 posts
Re: We are beginning to roll out new voice and image capabilities in ChatGPT
#642i am terrified now. at the rate this is going, i am sure it will plateau at somepoint, only thing that will stop/slow down progress is computation power.
'only thing that will stop/slow down progress is computation power'
Seems a bit contradictory? When has 'computation power' ever 'plateaued'?
Re: We are beginning to roll out new voice and image capabilities in ChatGPT
#643Re: We are beginning to roll out new voice and image capabilities in ChatGPT
#644Re: We are beginning to roll out new voice and image capabilities in ChatGPT
#645The number of comments here of people fearing there is a ghost in the shell is shocking. Are we really this emotional and irrational? Folks, let's all take a moment to remember that AI is nowhere near conscious. It's an illusion based in patterns that mimic humans.
Re: We are beginning to roll out new voice and image capabilities in ChatGPT
#646Earlier quoted context omitted.
I think this post-factual attitude is stronger and more common in some cultures than others. I'm afraid to say but given my extensive travels it appears American culture (and its derivatives in other countries) seems to be spearheading this shift.
Warning, my opinion ahead: I think it's because Americans, more than nearly all other cultures, love convenience. It's why the love for driving is so strong in the US. Don't walk or ride, drive. Once I was walking back from the grocer in Florida with 4 shopping bags, and people pulled over and asked if my car had broken down and if I needed a ride, people were stunned...I was walking for exercise and for the environm…
Re: We are beginning to roll out new voice and image capabilities in ChatGPT
#647There are a lot of comments attempting to rationalize the value add or differentiation of humans synthesizing information and communicating it to others vs an llm based ai doing something similar. The fact that it’s so difficult to find a compelling difference is insightful in itself.
Re: We are beginning to roll out new voice and image capabilities in ChatGPT
#648Earlier quoted context omitted.
I've replaced my voice google assistant searches with the voice feature of the Bing app. It's a night and day difference. Bing voice is what I always expected from an AI companion of the future, it is just lacking commands -- setting tasks, home automation, etc.
I got sick of searching Google for in-game recipes for Disney Dreamlight because most of the results are a bunch of pointless text, and then finally the recipe hidden in it somewhere. I used Bing yesterday and it was able to parse out exactly what I wanted, and then give me idiot-proof steps to making the recipe in-game. (I didn't need the steps, but it gave me what I wanted up front, easily.) I tried it twice and it…
You mean these? Took me a few seconds to find, not sure how an LLM would make that easier. I guess the biggest benefit of LLM then is for people who don't know how to find stuff.
Re: We are beginning to roll out new voice and image capabilities in ChatGPT
#649The thought of my children being put to bed by a machine is horrifying. Then again, perhaps this is better than many kids have. Shudder.
Re: We are beginning to roll out new voice and image capabilities in ChatGPT
#650Earlier quoted context omitted.
This feels fairly naive, ignoring how much progress has happened over the (short) span of one year. This doesn't sound like that tough of a gap to close in another year (again, projecting based off recent progress).
What actually was the innovation in LLMs that produced the kind of AI we're seeing now? Is that innovation ongoing or did it happen, and now we're seeing the various optimizations of that innovation? Is voice and image integration with ChatGPT a whole new capability of LLMs or is the "product" here a clean and intuitive interface through which to use the already existent technology? The difference between GPT 3, 3.5,…
Past the introduction of the transformer in 2017, There is no big "innovation". It is just scale. Bigger models are better. The last 4 years can be summed up that simply.
>Is voice and image integration with ChatGPT a whole new capability of LLMs or is the "product" here a clean and intuitive interface through which to use the already existent technology?
What is existing technology here ? Open ai aren't doing anything so alien you couldn't guess at if you knew what you were doing but image training at the scale of GPT-4 is new and it's not even the cleanest way to do it. We still don't have a "trained from scratch" large scale multimodal LLM yet.
>The difference between GPT 3, 3.5, and 4 is substantially smaller than the difference between GPT 2 and GPT 3
Definitely not lol. The OG GPT-3 was pulling sub 50 on MMLU. Even benchmarks aside, there is a massive gap in utility between 3.5 and 4, never mind 3. 4 was finished training august 2022. It's only 2 years apart from 3.
>I don't think progress is linear here. Rather, it seems more likely that we made the leap about a year or so ago, and are currently in the process of applying that leap in many different ways. But the leap happened, and there isn't seemingly another one coming.
There was no special leap (in terms of theory and engineering). This is scale plainly laid out and there's more of it to go.
>and Sam Altman has directly said there are no plans for a GPT 5.
the same that sat on 4 for 8 months and said absolutely nothing about it ? Take anything altman says about new iterations with a grain of salt.