Live data from Hacker News

We are beginning to roll out new voice and image capabilities in ChatGPT

openai.com

291–300 of 914 posts

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#291
post #61

Earlier quoted context omitted.

Talking to Google and Siri has been positively frustrating this year. On long solo drives, I just want to have a conversation to learn about random things. I've been itching to "talk" to chatGPT and learn more (french | music theory | history | math | whatever) all summer. This should hit the spot!

I still don't understand how you can talk to something that doesn't provide factual information and just take it at face value? The other day I asked it about the place I live and it made up nonsense, I was trying to get it to help me with an essay and it was just wrong, it was telling me things about this region that weren't real. Do we just drive through a town, ask for a made up history about it and just be satisf…

This is a fairly perpetual discussion, but I'll go for another round:

I feel like using LLM today is like using search 15 years ago - you get a feel for getting results you want.

I'd never use chatGPT for anything that's even remotely obscure, controversial, or niche.

But through all my double-checking, I've had phenomenal success rate in getting useful, readable, valid responses to well-covered / documented topics such as introductory french, introductory music theory, well-covered & non-controversial history and science.

I'd love to see the example you experienced; if I ask chatGPT "tell me about Toronto, Canada", my expectation would be to get high accuracy. If I asked it "Was Hum, Croatia, part of the Istrian liberation movement in the seventies", I'd have far less confidence - it's a leading question, on a less covered topic, introducing inaccuracies in the prompt.

My point is - for a 3 hour drive to cottage, I'm OK with something that's only 95% accurate on easy topics! I'd get no better from my spouse or best friend if they made it on the same drive :). My life will not depend on it, I'll have an educationally good time and miles will pass faster :).

(also, these conversations always seem to end in suffocatingly self-righteous "I don't know how others can live in this post-fact free world of ignorance", but that has a LOT of assumptions and, ironically, non-factual bias in it as well)

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#292
post #220

I'd like to see them put speech recognition through their LLM as a post-processing step. I find it's fairly common for whisper to make small but obvious mistakes (for example a word which is complete nonsense in the context of the sentence) which could be easily corrected for a similar sounding word that fits into the wider context of the sentence. Is anyone doing this? Is there a reason it doesn't work as well as I'…

Do you mean use the LLM as a post-processing step within a ChatGPT conversation? Or generally (like as part of Whisper)? If it’s the former, I’ve found that ChatGPT is good at working around transcription errors. Regarding the latter, I agree, but it wouldn’t be hard to use the GPT API for that.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#293

Earlier quoted context omitted.

You need a model of yourself to game out future scenarios, and that model or model+game is probably consciousness or very closely related. Sure, it's not completely in control but if it's just a rationalization then it begs the question: why bother? Is it accidental? If it's just an accident, then what replaces it in the planning process and why isn't that thing consciousness?

It's fine if you think that the planning process is what causes subjective experiences to arise. That may well be the case. I'm saying if you don't believe that non human objects can have subjective experiences, and then use that to define the limits of the behaviour of that object, that's a fallacy.

In humans, there seems to be a match between the subjective experience of consciousness and a high level planning job that needs doing. Our current LLMs are bad at high level planning, and it seems reasonable to suppose that making them good at high level planning might make them conscious or vice versa.

Agreed, woo is silly, but I didn't read it as woo but rather as a postulation that consciousness is what does high level planning.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#294
post #81
post #29

This announcement seem to have killed so many startups that were trying to do multi-modal on top of ChatGPT. The way it's progressing with solving use cases with images and voice, not too far when it might be the 'one app to rule them all'. I can already see "Alexa/Siri/Google Home" replacement, "Google Image Search" replacement, ed-tech startups that were solving problems with AI using by taking a photo are also doo…

I don't think anybody following OpenAI's feature releases will be caught off guard by ChatGPT becoming multi-modal. The app already features voice input. That still translates voice into text before sending, but it works so well that you basically never need to check or correct anything. Rather, you might have already been asking yourself why it doesn't reply back with a voice already. And the ability ingest images w…

one of the original training sets for the BERT series is called 'BookCorpus', accumulated by regular grad students for Natural Language Processing science. Part of the content was specifically and exactly purposed to "align" movies and video with written text. That is partly why it contains several thousand teen romance novels and ordinary paperback-style story telling content. What else is in there? "inquiring minds want to know"

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#295
post #18

The thought of my children being put to bed by a machine is horrifying. Then again, perhaps this is better than many kids have. Shudder.

If I could harness the power of AI to outsource my tasks, reading bedtime stories to my kids would be the last thing on that list. That's cherished time. Those are lifelong memories. Those are the moments we are supposed to be striving to have more of. It saddens me to think of the amount of engineering work that went into creating that example while entirely missing the point. These are the moments we are supposed t…

The AI takes care of the bedtime stories, giving you more time for video games.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#296
post #264
post #214

Earlier quoted context omitted.

When dealing with a tech where people have credible reasons to believe it can be enormously harmful on every possible time scale, maybe it would behoove them to not rocket down the interstate at 200mph?

Thats not what the analogy means. 200mph refers to funding.

No it refers to them moving too fast to send out basic emails for feature updates, per this comment chain.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#297

Earlier quoted context omitted.

Talking to Google and Siri has been positively frustrating this year. On long solo drives, I just want to have a conversation to learn about random things. I've been itching to "talk" to chatGPT and learn more (french | music theory | history | math | whatever) all summer. This should hit the spot!

Voice assistants have always been a half complete product. They were shown off as a cool feature, then they were never integrated so they were useful. The two biggest features I want are for the voice assistants to read something for me, and to do something on google/Apple Maps hand free. Neither of these ever work. “Siri/ ok google add the next gas station on the route” or “take me to the Chinese restaurant in Hobok…

In the current world:

Me: “OK Google, take me to the Chinese restaurant in Hoboken”

Google Assistant: “Calling Jessica Hobkin”.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#298
post #61

Earlier quoted context omitted.

I still don't understand how you can talk to something that doesn't provide factual information and just take it at face value? The other day I asked it about the place I live and it made up nonsense, I was trying to get it to help me with an essay and it was just wrong, it was telling me things about this region that weren't real. Do we just drive through a town, ask for a made up history about it and just be satisf…

What LLMs have made me realize more than anything is that we just don't care that much the information we receive being completely factual. I have tried to use it many times to learn a topic, and my experience has been that it is either frustratingly vague or incorrect. It's not a tool that I can completely add to my workflow until it is reliable, but I seem to be the odd one out.

ChatGPT 3.5 is terrible on technical subjects IME. Phind is best for me rn. Hugging Chat (Llama) works quite well too.

They're only good on universal truths. An amalgam of laws from around the globe doesn't tell me what the law is in my country, for example.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#299
post #174

Earlier quoted context omitted.

Maybe there is a state somewhere between "total chaos and carnage" and "emails users when new features are enabled for their account". Such as "decided it wasn't an operational priority to email users when features were enabled for them".

Emailing users when a new feature is enabled for their account isn't even the kind of thing that would distract an existing very busy developer. You could literally hire an entirely new guy, give him instructions to build such an email system, and let him put the right triggers on the user account permissions database to send out the right emails at the right time. And then, when it's built, you can start adding more…

The peanut gallery could (and does) say this about 1000 little features that 1000 different people claim to be so necessary that OpenAI is incompetent for not having, yet none of those people agree that those other features are priorities.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#300

Earlier quoted context omitted.

I work as a ethical hacker, so I'm well aware of the phishing and impersonation possibilities. But the net positive is so, so much bigger for society that I'm sure we'll figure it out. And yes, in 20 years you can tell your kids that 'back in my day' support consisted of real people. But truthfully, as someone who worked on a ISP helpdesk it's much better for society if these people move on to more productive areas.

I don’t think we know how these net out. AFAICT the negative use cases are a lot more real than the positive ones. People like to just suppose that these will help discover drugs and design buildings and what not, but what we actually know they’re capable of doing is littering our information environment at massive scale.

[flagged]
Post reply on HN