Live data from Hacker News

We are beginning to roll out new voice and image capabilities in ChatGPT

openai.com

241–250 of 914 posts

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#242
post #18

The thought of my children being put to bed by a machine is horrifying. Then again, perhaps this is better than many kids have. Shudder.

If I could harness the power of AI to outsource my tasks, reading bedtime stories to my kids would be the last thing on that list. That's cherished time. Those are lifelong memories. Those are the moments we are supposed to be striving to have more of. It saddens me to think of the amount of engineering work that went into creating that example while entirely missing the point. These are the moments we are supposed t…

I agree. I worry my culture is truly losing sight of what’s good in life. I don’t mean that as in “I know what’s best and everyone’s doing it wrong”, because I fully acknowledge that I can’t know what’s best for others. Yet I watch my friends and family work hard at things they don’t claim to value, I watch them lose life to scrolling and tv and movies they don’t actually enjoy, and I watch them lament that they don’t see their friends as much as they’d like, they don’t have enough time at home, kids are so much work, etc.

We have major priority issues from what I can see. If we want to live our lives more but put an AI to work doing something we tend to claim we place very high in our value hierarchy, we’re effectively inviting death into life. We’re forfeiting something we love. That’s incredibly sad to me.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#244
post #29

This announcement seem to have killed so many startups that were trying to do multi-modal on top of ChatGPT. The way it's progressing with solving use cases with images and voice, not too far when it might be the 'one app to rule them all'. I can already see "Alexa/Siri/Google Home" replacement, "Google Image Search" replacement, ed-tech startups that were solving problems with AI using by taking a photo are also doo…

It increasingly feels to me like building any kind of general-use AI tool or app is a bad choice. I see two viable AI business models:

1. Domain-specific AI - Training an AI model on highly technical and specific topics that general-purpose AI models don't excel at.

2. Integration - If you're going to build on an existing AI model, don't focus on adding more capabilities. Instead, focus on integrating it into companies' and users' existing workflows. Use it to automate internal processes and connect systems in ways that weren't previously possible. This adds a lot of value and isn't something that companies developing AI models are liable to do themselves.

The two will often go hand-in-hand.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#245

My biggest complaint with OpenAI/ChatGPT is their horrible "marketing" (for lack of a better term). They announce stuff like this (or like plugins), I get excited, I go to use it, it hasn't rolled out to me yet (which is frustrating as a paying customer), and my only recourse is.... check back daily? They never send an email "Plugins are available for you!", "Voice chat is now enabled on your account!" and so often I…

They're focused on scaling to meet the current (overwhelming) demand. Given the 'if you build it, they will come' dynamic they're experiencing, any focus on marketing would be a waste of resources.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#246
post #106
post #46

Earlier quoted context omitted.

In retrospect, such startups should have been wary: they should have known that OpenAI had Whisper, and also that GPT-4 was designed with image modality. I wouldn't say that OpenAI "telegraphed" their intentions, but the very first strategic question should have been, "Why isn't OpenAI doing this already, and what do we do if they decide to start?"

>I wouldn't say that OpenAI "telegraphed" their intentions They did telegraph it, they showed the multimodal capabilities back in the GPT4 Developer Livestream[0] right before first releasing it. 0. https://youtu.be/outcGtbnMuQ?t=943

Yeah I remember watching that and thinking oh I know a cool app idea. What if you just take a video of what food is in your kitchen and Chat GPT will create a recipe for you. I go to the docs and that was literally the example they gave.

I think the only place where plugins will make sense are for realtime things like booking travel or searching for sports/stock market/etc type information.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#247

Yet it still can't tell me how to import the Redirect type from Next.js and lies about it.

I don't know Next.js, but was that feature introduced later than 2021? I think both GPT-3.5 Turbo and GPT-4 largely share their datasets, and it has the data cutoff at roughly September 2021 (with a small amount of newer knowledge). This is their biggest drawback as of now to, say, Claude, which has a much newer dataset of early 2023.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#248
post #29

This announcement seem to have killed so many startups that were trying to do multi-modal on top of ChatGPT. The way it's progressing with solving use cases with images and voice, not too far when it might be the 'one app to rule them all'. I can already see "Alexa/Siri/Google Home" replacement, "Google Image Search" replacement, ed-tech startups that were solving problems with AI using by taking a photo are also doo…

Talking to Google and Siri has been positively frustrating this year. On long solo drives, I just want to have a conversation to learn about random things. I've been itching to "talk" to chatGPT and learn more (french | music theory | history | math | whatever) all summer. This should hit the spot!

Agreed. After using ChatGPT at all Siri is absolutely frustrating.

Example from a couple days ago:

Me, in the shower so not able to type: "Hey Siri, add 1.5 inch brad nails to my latest shopping list note."

Siri: "Sorry, I can't help with that."

... Really, Siri? You can't do something as simple as add a line to a note in the first-party Apple Notes app?

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#249

My biggest complaint with OpenAI/ChatGPT is their horrible "marketing" (for lack of a better term). They announce stuff like this (or like plugins), I get excited, I go to use it, it hasn't rolled out to me yet (which is frustrating as a paying customer), and my only recourse is.... check back daily? They never send an email "Plugins are available for you!", "Voice chat is now enabled on your account!" and so often I…

[deleted]

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#250

This is the dagger that will make online schooling unviable. ChatGPT already made it so that you could easily copy & paste any full-text questions and receive an answer with 90% accuracy. The only flaw was that problems that also used diagrams or figures would be out of the domain of ChatGPT. With image support, students could just take screenshots or document scans and have ChatGPT give them a valid answer. From wha…

When we talk about people abusing ChatGPT in a school context, it’s always for kids in high school or greater education levels. These are individuals that know right from wrong and also have the motor skills and access to use such a tool. These are individuals who are problem-solving for their specific need, which is to get this homework or essay out of the way so that they can do XYZ. Presumably XYZ does not leverage chatgpt. So make that what they spend their time on. At some point they’ll have to back-solve for skills they need to learn and need educational guidance and structure.

This is obviously not easy or going to happen without time and resources, but that is how adaptation goes.

Post reply on HN