Live data from Hacker News

We are beginning to roll out new voice and image capabilities in ChatGPT

openai.com

21–30 of 914 posts

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#22
post #4

The paper around GPT-4V(ision) which this uses: [0] Again. Model architecture and information is closed, as expected. [0] https://cdn.openai.com/papers/GPTV_System_Card.pdf

I wouldn’t call this a „paper“. They are pretty silent on a lot of technical details.

It's just a whitepaper.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#23
I just don't understand how they can package all of this for $20/m. Is compute really that cheap at scale?

I also wonder how Apple (& Google) is going be able to provide this for free? I would love to be fly in the meetings they have about this, imagine all the innovators dilemma like discussions they'd be forced to have (we have to do this vs this will eat up our margins).

This might be a little out there but I think Apple is making the correct move in letting the dust settle. Similar to how Zuckerberg burned $20 billion dollars for Apple to come out with Vision Pro, I see something similar playing out with Llama. Although this a low conviction take because software is Facebooks ballgame (hardware not so much).

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#24
post #9

Earlier quoted context omitted.

Its truly an amazing time to be alive. I'm right there with you, super excited about this decade. Especially what we could do in medicine.

Statistical diagnoses models have offered similar possibilities in medicine for 50 years. Pretty much, the idea is that you can get a far more accurate diagnosis if you take into account the medical history of everyone else in your family, town, workplace, residence and put all of it into a big statistical model, on top of your symptoms and history. However, medical secrecy, processes and laws prevent such things, ev…

See the glas half full or half empty?

Medical secrecy, processes and laws have indeed prevented SOME things, but a lot of things have gotten significantly better due to enhanced statistical models that have been implemented and widely used in real life scenarios.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#25
post #2

Old hat. This was done in 2009. ;) https://en.m.wikipedia.org/wiki/Project_Milo Milo had an AI structure that responded to human interactions, such as spoken word, gestures, or predefined actions in dynamic situations. The game relied on a procedural generation system which was constantly updating a built-in "dictionary" that was capable of matching key words in conversations with inherent voice-acting clips to simul…

Then Demis Hassabis ( Deepmind CEO ) probably worked on the tech while he was at LionHead as lead AI programmer on B&W.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#26

We should be fine as long as it doesn't move. Jokes aside, I have paused my subscription because even GPT4 seemed to become dumber at tasks to the point that I barely used it, but the constant influx of new features is tempting me to renew it just to check them out...

I read this all the time and yet no one can seem to come up with even a few questions from several months ago that ChatGPT has become “worse” at. You would think if this is happening it would be very easy to produce such evidence since chat history of all conversations is stored by default.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#28
I wonder how multimodal input and output will work with the chat API endpoints. I assume the messages array will contain URLs to an image, or maybe base64 encoded image data or something.

Maybe it will not be called the Chat API but rather the Multimodal API.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#29
This announcement seem to have killed so many startups that were trying to do multi-modal on top of ChatGPT. The way it's progressing with solving use cases with images and voice, not too far when it might be the 'one app to rule them all'.

I can already see "Alexa/Siri/Google Home" replacement, "Google Image Search" replacement, ed-tech startups that were solving problems with AI using by taking a photo are also doomed and more to follow.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#30

We should be fine as long as it doesn't move. Jokes aside, I have paused my subscription because even GPT4 seemed to become dumber at tasks to the point that I barely used it, but the constant influx of new features is tempting me to renew it just to check them out...

Did she/he said things like "I know I’ve made some very poor decisions recently, but I can give you my complete assurance that my work will be back to normal"?
Post reply on HN