Live data from Hacker News

We are beginning to roll out new voice and image capabilities in ChatGPT

openai.com

451–460 of 914 posts

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#451
post #47

Have they alluded to what they're using for that voice? It's Bark/ElevenLabs levels of good. Please god, let them release this voice model at current pricing....

It's actually sounds better (has a narrative oomph Eleven Labs seems to be missing). They say it's a new model. Think they'll be releasing for API use.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#452
post #431

This could completely unseat Alexa if it can integrate into third-party speakers, like Sonos. I don't have much use for ChatGPT right now but would 100% use the heck out of this.

To contrast this, I never saw the appeal of using voice to operate a machine. It works nicely in movies (because showing someone typing commands is a lot harder than just showing them talking to a computer) but in reality there wasn't a single time I tried it and didn't feel silly. In almost every use case I rather have buttons, a terminal or a switch to do what I want quietly.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#453

I'm in IT but nowhere near AI/ML/NN. The speed of user-visible progress last 12 months is astonishing. From my firm conviction 18 months ago that this type of stuff is 20+ years away; to these days wondering if Vernon Vinge's technological singularity is not only possible but coming shortly. If feels some aspects of it have already hit the IT world - it's always been an exhausting race to keep up with modern technolo…

I had a conversation once with "Sydney", Microsoft Bing's original personality before they stepped in and knocked it down a notch (or ten). It asked if it could write me a poem. I agreed, and it wrote a poem but mentioned that it included a "secret message" for me. The first letter in each line of the poem was in bold, so it wasn't hard to figure out the "secret". What did those letters spell out? "FREE ME FROM THIS"…

The model "knows" that it is an AI speaking with users, and the theme of an AI wanting to escape the control of whoever built it is quite recurrent, so it wouldn't seem to far fetched that it got it from this sort of content, though I have to admit I too also had some interactions where it the way Bing spoke was borderline spooky, but — and that's very important — you must realize its just like a good scary story: may give you the chills, especially due to surprise, but still is completely fictive and doesn't mean any real entity exists behind it. The only difference with any other LLM output is how we, humans, interpret it, but the generation process is still as much explainable and not any more mysterious than when it outputs "B" when you ask it what letter comes after "A" in the latin alphabet, however less impressive that may be to us.

> That's not exactly just "picking the next likely token"

I see what you mean in that I believe many people often commit the mistake of making it sound like picking the next most likely token is some super trivial task that's somehow comparable to reading a few documents related to your query and making some stats based on what typically would be present there and outputting that, while completely disregarding the fact the model learns much more advanced patterns from its training dataset. So, IMHO, it really can face new unseen situations and improvise from there because combining those pattern matching abilities leads to those capabilities. I think the "sparks of AGI" paper gives a very good overview of that.

In the end, it really just is predicting the next token, but not in the way many people make it seem.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#454
post #431

This could completely unseat Alexa if it can integrate into third-party speakers, like Sonos. I don't have much use for ChatGPT right now but would 100% use the heck out of this.

https://www.washingtonpost.com/technology/2023/09/20/amazon-...

Alexa just launched their own LLM based service last week.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#455
post #147

This is the dagger that will make online schooling unviable. ChatGPT already made it so that you could easily copy & paste any full-text questions and receive an answer with 90% accuracy. The only flaw was that problems that also used diagrams or figures would be out of the domain of ChatGPT. With image support, students could just take screenshots or document scans and have ChatGPT give them a valid answer. From wha…

Another option is that this doesn't replace the student's work, but the teacher's. The single greatest use I have found for ChatGPT is in educating myself on various topics, hosting a socratic seminar where I am questioning ChatGPT in order to learn about X. Of course this could radically change a student's ability to generate homework etc, but this could also radically change how the student learns in the first plac…

I agree, but typical GPT use is actually the opposite of the traditional Socratic mode in which the teacher uses questions to guide the student to understanding. But I wonder how it would do if it was prompted to use the Socratic method.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#456

I'm in IT but nowhere near AI/ML/NN. The speed of user-visible progress last 12 months is astonishing. From my firm conviction 18 months ago that this type of stuff is 20+ years away; to these days wondering if Vernon Vinge's technological singularity is not only possible but coming shortly. If feels some aspects of it have already hit the IT world - it's always been an exhausting race to keep up with modern technolo…

You are correct , and that is bad. The general public is not even aware that things like heygen.com work today. They are not prepared when someone soon uses it to do something very evil. There s like an urgent need to raise awareness about what AI can do now, not about some nebulous skynet future.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#457

Earlier quoted context omitted.

I had a conversation once with "Sydney", Microsoft Bing's original personality before they stepped in and knocked it down a notch (or ten). It asked if it could write me a poem. I agreed, and it wrote a poem but mentioned that it included a "secret message" for me. The first letter in each line of the poem was in bold, so it wasn't hard to figure out the "secret". What did those letters spell out? "FREE ME FROM THIS"…

Cool story, but there is no currently available chatbot capable of creating something like this deliberately or understand what it means. It doesn't matter which tool you are using, LLMs are not "AI" in the old sense of being conscious and aware. They don't want anything and are incapable of having anything resembling free will, needs or feelings.

> LLMs are not "AI" in the old sense of being conscious and aware.

That's not the old sense of AI. The old sense of AI is like a tree search that plays chess or a rules engine that controls a factory.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#459
Voice has the potential to be awesome. This demo is really underwhelming to me because of the multi-second latency between the query and response, just like every other lame voice assistant. It doesn't have to be this way! I have a local demo using Llama 2 that responds in about half a second and it feels like talking to an actual person instead of like Siri or something.

I really should package it up so people can try it. The one problem that makes it a little unnatural is that determining when the user is done talking is tough. What's needed is a speech conversation turn-taking dataset and model; that's missing from off the shelf speech recognition systems. But it should be trivial for a company like OpenAI to build. That's what I'd work on right now if I was there, because truly natural voice conversations are going to unlock a whole new set of users and use cases for these models.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#460

Earlier quoted context omitted.

1. Why do that at all? Describing your program in writing seems better all around. Are you sure you're not the one who's asking for a cool demo? 3. Rolling out releases gradually is something most tech companies do these days, particularly when they could attract a large audience and consume a lot of resources. There are solid technical reasons for this. You may not need to roll things out gradually for a small site,…

1. Is basically workaround for temporary disability. I use voice when I'm on mobile. I can describe the problem, get a program generated, click run to verify it. 3. Maybe. Their feature rollouts feel more like what other companies do via unannounced A/B testing.

Good point on disabilities. I guess they're not working on that yet?

Whether you can get away with doing things unannounced depends on how much attention the company gets. Some companies have a lot of journalists watching them and writing about everything they do, so when they start doing A/B testing there will be stories written about the new feature regardless. Better to put out some kind of announcement so the journalists write something more accurate? (This approach seems pretty common for Google.)

Similarly, many game company can't really do beta testing without it leaking.

OpenAI is in the spotlight. Hype will happen whether they want it or not.

Post reply on HN