Live data from Hacker News

We are beginning to roll out new voice and image capabilities in ChatGPT

openai.com

791–800 of 914 posts

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#791
As someone deep in the software test automation space, the thing I'm waiting for is robust AI-powered image recognition of app user interfaces. Combined with an AI ability to write test automation code, I'm looking forward to the ability to generate executable Selenium or Appium test code from a single screenshot (or sequence of screenshots). Feels like we're almost there.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#792
post #100

Okay the bike example is cute and impressive, but the human interaction seems to be obfuscating the potentially bigger application. With a few tweaks this is a general purpose solver for robotics planning. There are still a few hard problems between this and a working solution, but it is one of hard problems solved. Will we be seeing general purpose robots performing simple labor powered by chatgpt within the next ha…

I feel they could have used a more convincing example to be honest. Yeah it's cool it recognises so much but how useful is the demo in reality?

You have someone with a tool box and a manual (seriously who has a manual for their bike), asking the most basic question on how to lower a seatpost. My 5 year old kid knows how to do that.

Surely there's a better way to demonstrate the ground breaking impacts of ai on humanity than this. I dunno, something like how do I tie my shoelace.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#793

Earlier quoted context omitted.

now that you mention it, a big "for entertainment purposes only" banner like they use to have on all the psychic commercials on tv would not be inappropriate. it's incredible that LLMs are being marketed as general purpose assistants with a tiny asterisk, "may contain inaccuracies" like it's a walnut contamination

Not sure what's being incredible here. GPT-4 is a stellar general-purpose assistant, that shines when you stop treating it as encyclopedia, and start using it as an assistant . That is, give it tasks, like summarizing, or writing code, or explaining code, or rewriting prose. Ask for suggestions, ideas. You can do that to great effect, even when your requests are underspecified and somewhat confused, and it still work…

I just wish they were advertised for generative tasks and not retrieval tasks. It's not intelligence, it's not reasoning, it's text transformation.

It seems to be able to speak on history, sometimes it's even right, so there's a use case that people expect from it.

FYI I've used GPT4 and Claude 2 for hundreds of conversations, I understand what its good and bad at; I don't trust that the general public is being given a realistic view.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#794

Voice has the potential to be awesome. This demo is really underwhelming to me because of the multi-second latency between the query and response, just like every other lame voice assistant. It doesn't have to be this way! I have a local demo using Llama 2 that responds in about half a second and it feels like talking to an actual person instead of like Siri or something. I really should package it up so people can t…

Also curious to hear about your setup. Using whisper too? When I was experimenting with it there was still a lot of annoyance about hallucinations and I was hard coding some "if last phrase is 'thanks for watching', ignore last phrase" I was just googling a bit to see what's out there now for whisper/llama combos and came across this: https://github.com/yacineMTB/talk There's a demo linked on the github page that see…

Lol yeah the hallucinations are a huge problem. Likely solvable, I think there are probably some bugs in various whisper implementations that are making the problem worse than it should be. I haven't really dug in on that yet though. I was hoping I could switch to a different STT model more designed for real time like Meta's SeamlessM4T but it's still under a non-commercial license and I did have an idea that I might want to try making a product sometime. I did see that yacine made that version but I haven't tried it so I don't know how it compares to mine.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#795
post #574

Earlier quoted context omitted.

My unspoken thought-objects are wordless concepts, sounds, and images, with words only loosely hanging off those thought-objects. It takes additional effort to serialize thought-objects to sequences of words, and this is a lossy process - which would not be the case if I were thinking essentially in language.

You have no clue how GPT-4 functions so I don't know why you're assuming they're "thinking in language"

I am comfortable asserting that an LLM like GPT-4 is only capable of thinking in language; there is no distinction for an LLM between what it can conceive of and what it can express.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#796

Voice has the potential to be awesome. This demo is really underwhelming to me because of the multi-second latency between the query and response, just like every other lame voice assistant. It doesn't have to be this way! I have a local demo using Llama 2 that responds in about half a second and it feels like talking to an actual person instead of like Siri or something. I really should package it up so people can t…

> determining when the user is done talking is tough. Sometimes that task is tough for the speaker too, not just the listener. Courteous interruptions or the lack thereof might be a shibboleth for determining when we are speaking to an AI.

Yes interruptions are key, both ways. Having the user interrupt the bot is easy, but to have the bot interrupt the human will again require a model to predict when that should happen. But I do believe it is desirable for natural conversation.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#797

Voice has the potential to be awesome. This demo is really underwhelming to me because of the multi-second latency between the query and response, just like every other lame voice assistant. It doesn't have to be this way! I have a local demo using Llama 2 that responds in about half a second and it feels like talking to an actual person instead of like Siri or something. I really should package it up so people can t…

I wonder when computers will start taking our intonation into account too. That would really help with understanding the end of a phrase. And there’s SO MUCH information in intonation that doesn’t exist in pure text. Any AI that doesn’t understand that part of language will always still be kinda dumb, however clever they are.

You're right. Ultimately the only way this will really work is as an end-to-end model. Text will only get you so far. We could approximate it now with screenplay-like emotion annotations on text, which LLMs should both easily understand and be able to produce themselves (though you'd have to train a new speech recognition system to produce them). But end-to-end will be required eventually to reach human level fluency.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#798

Voice has the potential to be awesome. This demo is really underwhelming to me because of the multi-second latency between the query and response, just like every other lame voice assistant. It doesn't have to be this way! I have a local demo using Llama 2 that responds in about half a second and it feels like talking to an actual person instead of like Siri or something. I really should package it up so people can t…

> determining when the user is done talking is tough. Sometimes that task is tough for the speaker too, not just the listener. Courteous interruptions or the lack thereof might be a shibboleth for determining when we are speaking to an AI.

From prior experience, courteous interruption is a skill that a lot of humans find challenging at times too (myself included).

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#799

Earlier quoted context omitted.

Talking to Google and Siri has been positively frustrating this year. On long solo drives, I just want to have a conversation to learn about random things. I've been itching to "talk" to chatGPT and learn more (french | music theory | history | math | whatever) all summer. This should hit the spot!

I've wanted a ChatGPT Pod equivalent to a Google Home pod for a while! I have been intending to build it at some point. I am with you, talking to Google sucks. "Hey Google, why do ____ happen?" "I'm sorry, I don't know anything about that" But you're GOOGLE! Google it! What the heck lol So yeah, ChatGPT being able to hear what I say and give me info about it would be great! My holdup has been wakewords.

We have hardware and wake words:

https://heywillow.io/

Our REST endpoint can talk to whatever you want and we’ll have native ChatGPT soon.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#800

Voice has the potential to be awesome. This demo is really underwhelming to me because of the multi-second latency between the query and response, just like every other lame voice assistant. It doesn't have to be this way! I have a local demo using Llama 2 that responds in about half a second and it feels like talking to an actual person instead of like Siri or something. I really should package it up so people can t…

all it has to do is add a random selection of "uhms" and "ahhs" and "mmm"

Actually I do think this is a good idea. For best latency there should be multiple LLMs involved, a fast one to generate the first few words and then GPT-4 or similar for the rest of the response. In the case that the fast model is unsure, it could absolutely generate filler words while it waits for the big model to return the actual answer. I guess that's pretty much how humans use filler words too!
Post reply on HN