Live data from Hacker News

We are beginning to roll out new voice and image capabilities in ChatGPT

openai.com

111–120 of 914 posts

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#112
post #27

The picture feature would be amazing for tutorials. I can already imagine sending a photo of a synthesiser and asking ChatGPT to "turn the knobs" to make AI-generated presets

Man you're a genius. I was trying that uploading pdfs with manual of my synth and other stuff. With image that could be super easy.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#113
post #29

This announcement seem to have killed so many startups that were trying to do multi-modal on top of ChatGPT. The way it's progressing with solving use cases with images and voice, not too far when it might be the 'one app to rule them all'. I can already see "Alexa/Siri/Google Home" replacement, "Google Image Search" replacement, ed-tech startups that were solving problems with AI using by taking a photo are also doo…

It already replaced search engines. So much easier to write the question and explore the answers until it is solved.

This is funny, because I find it much less cumbersome to type a few search terms into a search engine and explore the links it spits out.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#114
post #9

Earlier quoted context omitted.

Its truly an amazing time to be alive. I'm right there with you, super excited about this decade. Especially what we could do in medicine.

Statistical diagnoses models have offered similar possibilities in medicine for 50 years. Pretty much, the idea is that you can get a far more accurate diagnosis if you take into account the medical history of everyone else in your family, town, workplace, residence and put all of it into a big statistical model, on top of your symptoms and history. However, medical secrecy, processes and laws prevent such things, ev…

The great thing about AI models is that once you train it, you can pretend the data wasn't illegal

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#115
post #61

Earlier quoted context omitted.

Talking to Google and Siri has been positively frustrating this year. On long solo drives, I just want to have a conversation to learn about random things. I've been itching to "talk" to chatGPT and learn more (french | music theory | history | math | whatever) all summer. This should hit the spot!

I still don't understand how you can talk to something that doesn't provide factual information and just take it at face value? The other day I asked it about the place I live and it made up nonsense, I was trying to get it to help me with an essay and it was just wrong, it was telling me things about this region that weren't real. Do we just drive through a town, ask for a made up history about it and just be satisf…

Because it doesn't always make up stuff. Because I'm a human and can ask for more information. I don't want an encyclopedia on a podcast. I want to "talk" to someone about stuff. Not have an enumerated list of truths firehosed at me.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#116
post #49

I'm in IT but nowhere near AI/ML/NN. The speed of user-visible progress last 12 months is astonishing. From my firm conviction 18 months ago that this type of stuff is 20+ years away; to these days wondering if Vernon Vinge's technological singularity is not only possible but coming shortly. If feels some aspects of it have already hit the IT world - it's always been an exhausting race to keep up with modern technolo…

I also don't believe LLMs are "conscious", but I also don't know what that means, and I have yet to see a definition of "statistically guessing next word" that cannot be applied to what a human brain does to generate the next word.

Your brain doesn't solely pick the next best word. As best as I understand it, the brain has an external state of the world that constantly updates, paired to an internal model predicting the next best word.

Which is why we can create the counterfactual that "The Cowboys should have won last night" and it has implicit meaning.

Current LLM models don't have an external state of the world, which is why folks like LeCunn are suggesting model architectures like JEPA. Without an external, correcting state of the world, model prediction errors compound almost surely (to use a technical phrase).

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#117
post #49

Earlier quoted context omitted.

I also don't believe LLMs are "conscious", but I also don't know what that means, and I have yet to see a definition of "statistically guessing next word" that cannot be applied to what a human brain does to generate the next word.

When my brain generates the next wurd I'm perfectly capable of taking decisions of misspelling "word" for "wurd", LLMs can't make such reasonings unless instructed to act like that.

Why would I want an AI assistant to have agency? I want them to help me, not to further their personal goals. In fact, I don't want them to have personal goals other than helping people.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#118
post #61

Earlier quoted context omitted.

Talking to Google and Siri has been positively frustrating this year. On long solo drives, I just want to have a conversation to learn about random things. I've been itching to "talk" to chatGPT and learn more (french | music theory | history | math | whatever) all summer. This should hit the spot!

I still don't understand how you can talk to something that doesn't provide factual information and just take it at face value? The other day I asked it about the place I live and it made up nonsense, I was trying to get it to help me with an essay and it was just wrong, it was telling me things about this region that weren't real. Do we just drive through a town, ask for a made up history about it and just be satisf…

I'm curious if you're using GPT-4 ($)? I find a lot of the criticisms about hallucination come from users who aren't, and my experience with GPT-4 is it's far less likely to make stuff up. Does it know all the answers, certainly not, but it's self-aware enough to say sorry I don't know instead of making a wild guess.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#119
post #90
post #79

Earlier quoted context omitted.

Agreed except ChatGPT (3.5 at least, haven't tried 4) is unable to provide primary sources for its results. At least when I tried, it just provided hallucinated urls

Try it. There's a world of difference.

In general or for this specific application (linking primary sources)?

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#120

I'm in IT but nowhere near AI/ML/NN. The speed of user-visible progress last 12 months is astonishing. From my firm conviction 18 months ago that this type of stuff is 20+ years away; to these days wondering if Vernon Vinge's technological singularity is not only possible but coming shortly. If feels some aspects of it have already hit the IT world - it's always been an exhausting race to keep up with modern technolo…

One cool aspect of LLMs is Vernon Vinge's programming archaeology needn't be a thing... LLMs can go down every code path and identify what it does, when it was added, and whether it's still needed.

It might even be correct. Occasionally.
Post reply on HN