Live data from Hacker News

We are beginning to roll out new voice and image capabilities in ChatGPT

openai.com

71–80 of 914 posts

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#71

We should be fine as long as it doesn't move. Jokes aside, I have paused my subscription because even GPT4 seemed to become dumber at tasks to the point that I barely used it, but the constant influx of new features is tempting me to renew it just to check them out...

For me the most glaring example of this was it's document parsong capability in GPT4. I was using it to revamp my resume. I would upload it to got, ask for suggestions, incorporate them into the word document and then repeat the steps till I was satisfied.

After maybe 3 iterations gpt4 started claiming that it is not capable of reading from a word document even though it's done that the last 3 times. Have to click regenerate button to get it to work

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#72
post #49

I'm in IT but nowhere near AI/ML/NN. The speed of user-visible progress last 12 months is astonishing. From my firm conviction 18 months ago that this type of stuff is 20+ years away; to these days wondering if Vernon Vinge's technological singularity is not only possible but coming shortly. If feels some aspects of it have already hit the IT world - it's always been an exhausting race to keep up with modern technolo…

I also don't believe LLMs are "conscious", but I also don't know what that means, and I have yet to see a definition of "statistically guessing next word" that cannot be applied to what a human brain does to generate the next word.

When my brain generates the next wurd I'm perfectly capable of taking decisions of misspelling "word" for "wurd", LLMs can't make such reasonings unless instructed to act like that.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#73
post #61

Earlier quoted context omitted.

Talking to Google and Siri has been positively frustrating this year. On long solo drives, I just want to have a conversation to learn about random things. I've been itching to "talk" to chatGPT and learn more (french | music theory | history | math | whatever) all summer. This should hit the spot!

I still don't understand how you can talk to something that doesn't provide factual information and just take it at face value? The other day I asked it about the place I live and it made up nonsense, I was trying to get it to help me with an essay and it was just wrong, it was telling me things about this region that weren't real. Do we just drive through a town, ask for a made up history about it and just be satisf…

What LLMs have made me realize more than anything is that we just don't care that much the information we receive being completely factual.

I have tried to use it many times to learn a topic, and my experience has been that it is either frustratingly vague or incorrect.

It's not a tool that I can completely add to my workflow until it is reliable, but I seem to be the odd one out.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#74
post #28

I wonder how multimodal input and output will work with the chat API endpoints. I assume the messages array will contain URLs to an image, or maybe base64 encoded image data or something. Maybe it will not be called the Chat API but rather the Multimodal API.

Are there already some rumors on when the multimodal API will be available?

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#75
post #61

Earlier quoted context omitted.

Talking to Google and Siri has been positively frustrating this year. On long solo drives, I just want to have a conversation to learn about random things. I've been itching to "talk" to chatGPT and learn more (french | music theory | history | math | whatever) all summer. This should hit the spot!

I still don't understand how you can talk to something that doesn't provide factual information and just take it at face value? The other day I asked it about the place I live and it made up nonsense, I was trying to get it to help me with an essay and it was just wrong, it was telling me things about this region that weren't real. Do we just drive through a town, ask for a made up history about it and just be satisf…

> how you can talk to something that doesn't provide factual information and just take it at face value

Like talking to most people you mean?

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#76
post #29

This announcement seem to have killed so many startups that were trying to do multi-modal on top of ChatGPT. The way it's progressing with solving use cases with images and voice, not too far when it might be the 'one app to rule them all'. I can already see "Alexa/Siri/Google Home" replacement, "Google Image Search" replacement, ed-tech startups that were solving problems with AI using by taking a photo are also doo…

This is good news - those ai companies have been freed to work on something else, along with the ai workers they employ. This is of great benefit to society.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#77
post #29

This announcement seem to have killed so many startups that were trying to do multi-modal on top of ChatGPT. The way it's progressing with solving use cases with images and voice, not too far when it might be the 'one app to rule them all'. I can already see "Alexa/Siri/Google Home" replacement, "Google Image Search" replacement, ed-tech startups that were solving problems with AI using by taking a photo are also doo…

"Don't build your castle in someone else's kingdom."

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#79
post #29

This announcement seem to have killed so many startups that were trying to do multi-modal on top of ChatGPT. The way it's progressing with solving use cases with images and voice, not too far when it might be the 'one app to rule them all'. I can already see "Alexa/Siri/Google Home" replacement, "Google Image Search" replacement, ed-tech startups that were solving problems with AI using by taking a photo are also doo…

It already replaced search engines. So much easier to write the question and explore the answers until it is solved.

Agreed except ChatGPT (3.5 at least, haven't tried 4) is unable to provide primary sources for its results. At least when I tried, it just provided hallucinated urls

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#80
post #49

I'm in IT but nowhere near AI/ML/NN. The speed of user-visible progress last 12 months is astonishing. From my firm conviction 18 months ago that this type of stuff is 20+ years away; to these days wondering if Vernon Vinge's technological singularity is not only possible but coming shortly. If feels some aspects of it have already hit the IT world - it's always been an exhausting race to keep up with modern technolo…

I also don't believe LLMs are "conscious", but I also don't know what that means, and I have yet to see a definition of "statistically guessing next word" that cannot be applied to what a human brain does to generate the next word.

> I have yet to see a definition of "statistically guessing next word" that cannot be applied to what a human brain does to generate the next word.

I think this is true. The problem is equating this process with how humans think though.

Post reply on HN