Live data from Hacker News

We are beginning to roll out new voice and image capabilities in ChatGPT

openai.com

121–130 of 914 posts

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#121
post #45
post #18

The thought of my children being put to bed by a machine is horrifying. Then again, perhaps this is better than many kids have. Shudder.

I actually think that what is sad is that it seems as if having viable future as a creative visual artist is likely done. This was a major, major, major outlet and sanctuary for certain types of people to find meaning and fulfillment in their life which is now in the process of being wiped out for a quick buck. We'll be told by OpenAI and friends is that it shouldn't be a problem, because those were mundane tasks and…

fwiw the only piece of AI art that has given me the sense of awe and beauty that art you'd find in a museum gives me was that spiral town image https://twitter.com/MrUgleh/status/1705316060201681313, which is something you couldn't have really made without AI. But that was only interesting because of the unique human generated idea behind it which was the encoding of a geometric pattern within a scene.

Most AI art is just generic garbage that you scroll past immediately and doesn't offer you anything.

We're gonna have to do something to stop the biggest crisis in meaning ever that comes out of this eventually though. Eventually no one will be of any economic value to society. Maybe just put someone in an ultra realistic simulation to give them artificial meaning.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#122

Earlier quoted context omitted.

Talking to Google and Siri has been positively frustrating this year. On long solo drives, I just want to have a conversation to learn about random things. I've been itching to "talk" to chatGPT and learn more (french | music theory | history | math | whatever) all summer. This should hit the spot!

I've replaced my voice google assistant searches with the voice feature of the Bing app. It's a night and day difference. Bing voice is what I always expected from an AI companion of the future, it is just lacking commands -- setting tasks, home automation, etc.

precisely this. once someone figures out how to get something like GPT integrated with actual products like smart home devices and the same access levels as siri/google assistant, it will be the true voice assistant experience everyone has wanted.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#123

Earlier quoted context omitted.

Statistical diagnoses models have offered similar possibilities in medicine for 50 years. Pretty much, the idea is that you can get a far more accurate diagnosis if you take into account the medical history of everyone else in your family, town, workplace, residence and put all of it into a big statistical model, on top of your symptoms and history. However, medical secrecy, processes and laws prevent such things, ev…

This is what effectively doctors do - educated guessing. In my view, while statistical models would probably be an improvement ( assuming all confounding factors are measured ), the ultimate solution is not to get better at educated guessing, but to remove the guessing completely, with diagnostic tests that measure the relevant bio-medical markers.

Good tests This becomes even more true when you consider there is risk to every test. Some tests have obvious risks (radiation risk from CT scans, chance of damage from spinal fluid tap). Other tests the risk is less obvious (sending you for a blood test and awaiting the results might not be a good idea if that delays treatment for some ailment already pretty certain). In the bigger picture, any test that costs money harms the patient slightly, since someone must pay for the test, and for many the money they spend on extra tests comes out of money they might otherwise spend on gym memberships, better food, or working fewer hours - it is well known that the poor have worse health than the rich.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#124
post #29

This announcement seem to have killed so many startups that were trying to do multi-modal on top of ChatGPT. The way it's progressing with solving use cases with images and voice, not too far when it might be the 'one app to rule them all'. I can already see "Alexa/Siri/Google Home" replacement, "Google Image Search" replacement, ed-tech startups that were solving problems with AI using by taking a photo are also doo…

Talking to Google and Siri has been positively frustrating this year. On long solo drives, I just want to have a conversation to learn about random things. I've been itching to "talk" to chatGPT and learn more (french | music theory | history | math | whatever) all summer. This should hit the spot!

It’s funny. Driving buddy has been my number one use case for a while now.

Still can’t quite make it work. I feel like I could learn a lot if I could have random conversations with GPT.

+ bonus if someone else in the car got excited when I see cows. Don’t care if it’s an AI.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#125
post #89

Does anyone know how they linked image recognition with an LLM to give such specific instructions as shown in the bike video on the website?

I don't know but GPT4 was multimodal from the beginning. They just delayed the release of its image processing abilities.

> We’ve created GPT-4, the latest milestone in OpenAI’s effort in scaling up deep learning. GPT-4 is a large multimodal model (accepting image and text inputs, emitting text outputs) that, while less capable than humans in many real-world scenarios, exhibits human-level performance on various professional and academic benchmarks.

> March 14, 2023

https://openai.com/research/gpt-4

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#126
post #49

I'm in IT but nowhere near AI/ML/NN. The speed of user-visible progress last 12 months is astonishing. From my firm conviction 18 months ago that this type of stuff is 20+ years away; to these days wondering if Vernon Vinge's technological singularity is not only possible but coming shortly. If feels some aspects of it have already hit the IT world - it's always been an exhausting race to keep up with modern technolo…

I also don't believe LLMs are "conscious", but I also don't know what that means, and I have yet to see a definition of "statistically guessing next word" that cannot be applied to what a human brain does to generate the next word.

Another aspect is "is the output good enough for what it's meant to do?"

We don't need "originality" or "human creativity" - if a certain AI-generated piece of content does its job, it's "good enough".

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#127
post #49

I'm in IT but nowhere near AI/ML/NN. The speed of user-visible progress last 12 months is astonishing. From my firm conviction 18 months ago that this type of stuff is 20+ years away; to these days wondering if Vernon Vinge's technological singularity is not only possible but coming shortly. If feels some aspects of it have already hit the IT world - it's always been an exhausting race to keep up with modern technolo…

I also don't believe LLMs are "conscious", but I also don't know what that means, and I have yet to see a definition of "statistically guessing next word" that cannot be applied to what a human brain does to generate the next word.

I believe that the distinguishing factor between what an LLM and a human brain do to generate the next word is that the human brain expresses intentionality originating from inner states and future expectations. As I type this comment I'm sure one could argue that the biological neural networks in my brain are choosing the next word based on statistical guessing, and that the initial prompt was your initial comment.

What sets my brain apart from an LLM though is that I am not typing this because you asked me to do it, nor because I needed to reply to the first comment I saw. I am typing this because it is a thought that has been in my mind for a while and I am interested in expressing it to other human brains, motivated by a mix of arrogant belief that it is insightful and a wish to see others either agreeing or providing reasonable counterpoints—I have an intention behind it. And, equally relevant, I must make an effort to not elaborate any more on this point because I have the conflicting intention to leave my laptop and do other stuff.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#128
post #100

Okay the bike example is cute and impressive, but the human interaction seems to be obfuscating the potentially bigger application. With a few tweaks this is a general purpose solver for robotics planning. There are still a few hard problems between this and a working solution, but it is one of hard problems solved. Will we be seeing general purpose robots performing simple labor powered by chatgpt within the next ha…

> With a few tweaks this is a general purpose solver for robotics planning.

Yeah, but with an enormous ecological footprint.

Also, not suitable for small lightweight robots like drones.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#130
post #61

Earlier quoted context omitted.

Talking to Google and Siri has been positively frustrating this year. On long solo drives, I just want to have a conversation to learn about random things. I've been itching to "talk" to chatGPT and learn more (french | music theory | history | math | whatever) all summer. This should hit the spot!

I still don't understand how you can talk to something that doesn't provide factual information and just take it at face value? The other day I asked it about the place I live and it made up nonsense, I was trying to get it to help me with an essay and it was just wrong, it was telling me things about this region that weren't real. Do we just drive through a town, ask for a made up history about it and just be satisf…

> I still don't understand how you can talk to something that doesn't provide factual information and just take it at face value?

All human interactions from all of history called and they …

Post reply on HN