Live data from Hacker News

We are beginning to roll out new voice and image capabilities in ChatGPT

openai.com

341–350 of 914 posts

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#341
post #100

Okay the bike example is cute and impressive, but the human interaction seems to be obfuscating the potentially bigger application. With a few tweaks this is a general purpose solver for robotics planning. There are still a few hard problems between this and a working solution, but it is one of hard problems solved. Will we be seeing general purpose robots performing simple labor powered by chatgpt within the next ha…

That bike example seemed a mix of underwhelming (for being the demo video) and even confusing.

1. It's not smart enough to recognize from the initial image this is a bolt style seat lock (which a human can).

2. The manual is not shown to the viewer, so I can't infer how the model knows this is a 4mm bolt (or if it is just guessing given that's the most likely one).

3. I don't understand how it can know the toolbox is using metric allen wrenches.

Additionally is this just the same vision model that exists in bing chat?

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#342
post #18

The thought of my children being put to bed by a machine is horrifying. Then again, perhaps this is better than many kids have. Shudder.

If I could harness the power of AI to outsource my tasks, reading bedtime stories to my kids would be the last thing on that list. That's cherished time. Those are lifelong memories. Those are the moments we are supposed to be striving to have more of. It saddens me to think of the amount of engineering work that went into creating that example while entirely missing the point. These are the moments we are supposed t…

I remember in the "microsoft office Generative AI" demo, one of the motivating examples was a parent generating a graduation party speech for her child... [1]

The first half of the video is demonstrating how the parent can take something as special as a party celebrating a major milestone and automate it into a soulless box-check – while editing some segments to make it look like their own voice.

Definite black mirror vibes.

[1]: https://youtu.be/ebls5x-gb0s?t=224

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#343
post #105
post #74

Earlier quoted context omitted.

Are there already some rumors on when the multimodal API will be available?

The announcement says after the Plus rollout then it will go in the API.

Where does it say that?

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#344
post #325

Earlier quoted context omitted.

> My biggest complaint with OpenAI/ChatGPT is their horrible "marketing" Agreed. Other notable mentions: choosing "ChatGPT" as their product name and not having mobile apps.

They do have mobile apps though?

Oops, missed that announcement.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#345
post #18

The thought of my children being put to bed by a machine is horrifying. Then again, perhaps this is better than many kids have. Shudder.

If I could harness the power of AI to outsource my tasks, reading bedtime stories to my kids would be the last thing on that list. That's cherished time. Those are lifelong memories. Those are the moments we are supposed to be striving to have more of. It saddens me to think of the amount of engineering work that went into creating that example while entirely missing the point. These are the moments we are supposed t…

I viewed this differently. This wasn't a parent having an AI step in to read their kid a bedtime story, it was a parent and a child using AI to discover an interesting story together.

It's just like reading a "choose your own adventure" book with your child, but it can be much more interactive and you both come up with ideas and have the LLM integrate them.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#346
> The new voice capability is powered by a new text-to-speech model, capable of generating human-like audio from just text and a few seconds of sample speech.

Sadly, they lost the "open" since a long ago... Would be wonderful to have these models open sourced...

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#347

My biggest complaint with OpenAI/ChatGPT is their horrible "marketing" (for lack of a better term). They announce stuff like this (or like plugins), I get excited, I go to use it, it hasn't rolled out to me yet (which is frustrating as a paying customer), and my only recourse is.... check back daily? They never send an email "Plugins are available for you!", "Voice chat is now enabled on your account!" and so often I…

It has always seemed like OpenAI succeeds in spite of itself. API access was historically an absolute nightmare, and it just seemed like they didn't even want customers.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#348

Now just throw this into a humanoid looking robot with fine motor skills and we are halfway to a dystopian hellscape that is now only years away instead of decades. What a time to be alive.

What would make it dystopian would be if this humanoid robot was then granted rights. As a servant, it could be useful.

I would like our future Cylon overlords to know that I had nothing to do with this!

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#349

Earlier quoted context omitted.

They do explain why in the post. (Still, you may not agree, of course.) > We are deploying image and voice capabilities gradually > > OpenAI’s goal is to build AGI that is safe and beneficial. We believe in making our tools available gradually, which allows us to make improvements and refine risk mitigations over time while also preparing everyone for more powerful systems in the future. This strategy becomes even mo…

My issue isn't fully with them rolling out slowly, my issue is never knowing when you will get the feature or rather not being told when you do get it. I'm fine with "sometime in the next X days/months you will get feature Y", my issue is the only way to see if you got feature Y is to check back daily.

It's the first sentence in the 3rd paragraph, repeated again at the end of the blog post.

> We’re rolling out voice and images in ChatGPT to Plus and Enterprise users over the next two weeks.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#350

Earlier quoted context omitted.

Talking to Google and Siri has been positively frustrating this year. On long solo drives, I just want to have a conversation to learn about random things. I've been itching to "talk" to chatGPT and learn more (french | music theory | history | math | whatever) all summer. This should hit the spot!

It’s funny. Driving buddy has been my number one use case for a while now. Still can’t quite make it work. I feel like I could learn a lot if I could have random conversations with GPT. + bonus if someone else in the car got excited when I see cows. Don’t care if it’s an AI.

Try Pi AI. They have an app that can be voice/audio driven. Works well for the driving buddy scenario.

https://pi.ai/talk

Post reply on HN