Live data from Hacker News

We are beginning to roll out new voice and image capabilities in ChatGPT

openai.com

51–60 of 914 posts

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#51
post #29

This announcement seem to have killed so many startups that were trying to do multi-modal on top of ChatGPT. The way it's progressing with solving use cases with images and voice, not too far when it might be the 'one app to rule them all'. I can already see "Alexa/Siri/Google Home" replacement, "Google Image Search" replacement, ed-tech startups that were solving problems with AI using by taking a photo are also doo…

It already replaced search engines. So much easier to write the question and explore the answers until it is solved.

It rather created new hybrid search engines, like perplexity and phind.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#52
post #28

I wonder how multimodal input and output will work with the chat API endpoints. I assume the messages array will contain URLs to an image, or maybe base64 encoded image data or something. Maybe it will not be called the Chat API but rather the Multimodal API.

AIPI

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#53

Earlier quoted context omitted.

Statistical diagnoses models have offered similar possibilities in medicine for 50 years. Pretty much, the idea is that you can get a far more accurate diagnosis if you take into account the medical history of everyone else in your family, town, workplace, residence and put all of it into a big statistical model, on top of your symptoms and history. However, medical secrecy, processes and laws prevent such things, ev…

See the glas half full or half empty? Medical secrecy, processes and laws have indeed prevented SOME things, but a lot of things have gotten significantly better due to enhanced statistical models that have been implemented and widely used in real life scenarios.

To make this feasible (meaning that the TB of data and the huge computing effort is somewhere else, and I only have the mic (smartphone), we need our local agent to send multiple irrelevant queries to the mothership, to hide our true purpose.

Example: my favourite team is X. So if I want to keep it a secret, when I ask for the history of championships of X, I will ask for X. My local agent should ask for 100 teams, get all the data, and then report back for only X. Eventually the mothership will figure out what we like (a large wenn diagram). But this is not in anyone's interest, and thus will not happen.

Also, like this the local agent will be able to learn and remember us, at a cost.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#55
post #29

This announcement seem to have killed so many startups that were trying to do multi-modal on top of ChatGPT. The way it's progressing with solving use cases with images and voice, not too far when it might be the 'one app to rule them all'. I can already see "Alexa/Siri/Google Home" replacement, "Google Image Search" replacement, ed-tech startups that were solving problems with AI using by taking a photo are also doo…

Talking to Google and Siri has been positively frustrating this year. On long solo drives, I just want to have a conversation to learn about random things. I've been itching to "talk" to chatGPT and learn more (french | music theory | history | math | whatever) all summer. This should hit the spot!

I assume you have never heard of podcasts.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#56

I know there are shades of grey to how they operate, but the near constant stream of stuff they're shipping keeps me excited. The LLM boom of the last year (Open AI, llama, et al) has me giddy as a software person. It's a reach, but I truly feel like I'm watching the pyramids of our time get made.

Computers understanding and responding in human language is the most exciting software innovation since the invention of the GUI. Just as the GUI made computer software available to billions LLMs will be the next revolution. I'm just as excited as you! The only downside is that it now make me feel bad that I'm not doing anything with it yet.

> The only downside is that it now make me feel bad that I'm not doing anything with it yet.

If that's the only downside that you see... I guess enhanced phishing/impersonation and all the blackhat stuff that come with it don't count.

I for one already miss the time where companies had support teams made of actual people.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#57
post #9

Earlier quoted context omitted.

Its truly an amazing time to be alive. I'm right there with you, super excited about this decade. Especially what we could do in medicine.

Statistical diagnoses models have offered similar possibilities in medicine for 50 years. Pretty much, the idea is that you can get a far more accurate diagnosis if you take into account the medical history of everyone else in your family, town, workplace, residence and put all of it into a big statistical model, on top of your symptoms and history. However, medical secrecy, processes and laws prevent such things, ev…

This is what effectively doctors do - educated guessing.

In my view, while statistical models would probably be an improvement ( assuming all confounding factors are measured ), the ultimate solution is not to get better at educated guessing, but to remove the guessing completely, with diagnostic tests that measure the relevant bio-medical markers.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#58

We should be fine as long as it doesn't move. Jokes aside, I have paused my subscription because even GPT4 seemed to become dumber at tasks to the point that I barely used it, but the constant influx of new features is tempting me to renew it just to check them out...

I read this all the time and yet no one can seem to come up with even a few questions from several months ago that ChatGPT has become “worse” at. You would think if this is happening it would be very easy to produce such evidence since chat history of all conversations is stored by default.

OpenAI regularly changes the model and they admit the new models are more restricted, in the sense that they prevent tricky prompts from producing naughty words, etc.

It should be their responsibility to prove that it's just as capable.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#59

Earlier quoted context omitted.

Talking to Google and Siri has been positively frustrating this year. On long solo drives, I just want to have a conversation to learn about random things. I've been itching to "talk" to chatGPT and learn more (french | music theory | history | math | whatever) all summer. This should hit the spot!

I assume you have never heard of podcasts.

you can ask podcasts questions? and they answer you?

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#60
I'm following on trying to understand how close I am to developing my personal coding assistant I can speak with.

Doesn't really need to do much besides writing down my tasks/todos and updating them, occasionally maybe provide feedback or write a code snippet. This all seems in the current capabilities of OpenAI's offering.

Sadly voice chat is still not available on PC where I do my development.

Post reply on HN