This announcement seem to have killed so many startups that were trying to do multi-modal on top of ChatGPT. The way it's progressing with solving use cases with images and voice, not too far when it might be the 'one app to rule them all'. I can already see "Alexa/Siri/Google Home" replacement, "Google Image Search" replacement, ed-tech startups that were solving problems with AI using by taking a photo are also doo…
In retrospect, such startups should have been wary: they should have known that OpenAI had Whisper, and also that GPT-4 was designed with image modality. I wouldn't say that OpenAI "telegraphed" their intentions, but the very first strategic question should have been, "Why isn't OpenAI doing this already, and what do we do if they decide to start?"
We are beginning to roll out new voice and image capabilities in ChatGPT
101–110 of 914 posts
Re: We are beginning to roll out new voice and image capabilities in ChatGPT
#102This announcement seem to have killed so many startups that were trying to do multi-modal on top of ChatGPT. The way it's progressing with solving use cases with images and voice, not too far when it might be the 'one app to rule them all'. I can already see "Alexa/Siri/Google Home" replacement, "Google Image Search" replacement, ed-tech startups that were solving problems with AI using by taking a photo are also doo…
It already replaced search engines. So much easier to write the question and explore the answers until it is solved.
ChatGPT is my primary search engine now. (I just wish it would accept a URL query parameter so it could be launched straight from the browser address bar.)
Re: We are beginning to roll out new voice and image capabilities in ChatGPT
#103Earlier quoted context omitted.
Talking to Google and Siri has been positively frustrating this year. On long solo drives, I just want to have a conversation to learn about random things. I've been itching to "talk" to chatGPT and learn more (french | music theory | history | math | whatever) all summer. This should hit the spot!
I assume you have never heard of podcasts.
Re: We are beginning to roll out new voice and image capabilities in ChatGPT
#104They could also improve their current features. I always need to regenerate answers.
Re: We are beginning to roll out new voice and image capabilities in ChatGPT
#105I wonder how multimodal input and output will work with the chat API endpoints. I assume the messages array will contain URLs to an image, or maybe base64 encoded image data or something. Maybe it will not be called the Chat API but rather the Multimodal API.
Are there already some rumors on when the multimodal API will be available?
Re: We are beginning to roll out new voice and image capabilities in ChatGPT
#106This announcement seem to have killed so many startups that were trying to do multi-modal on top of ChatGPT. The way it's progressing with solving use cases with images and voice, not too far when it might be the 'one app to rule them all'. I can already see "Alexa/Siri/Google Home" replacement, "Google Image Search" replacement, ed-tech startups that were solving problems with AI using by taking a photo are also doo…
In retrospect, such startups should have been wary: they should have known that OpenAI had Whisper, and also that GPT-4 was designed with image modality. I wouldn't say that OpenAI "telegraphed" their intentions, but the very first strategic question should have been, "Why isn't OpenAI doing this already, and what do we do if they decide to start?"
They did telegraph it, they showed the multimodal capabilities back in the GPT4 Developer Livestream[0] right before first releasing it.
Re: We are beginning to roll out new voice and image capabilities in ChatGPT
#107Earlier quoted context omitted.
Its truly an amazing time to be alive. I'm right there with you, super excited about this decade. Especially what we could do in medicine.
Statistical diagnoses models have offered similar possibilities in medicine for 50 years. Pretty much, the idea is that you can get a far more accurate diagnosis if you take into account the medical history of everyone else in your family, town, workplace, residence and put all of it into a big statistical model, on top of your symptoms and history. However, medical secrecy, processes and laws prevent such things, ev…
> Similar possibilities existed in medicine for 50 years
It would've been like building the tower of babel with a bunch of raspbery pi zeros. While theoretically possible, practically impossible and not (just) because of laws, but rather because of structural limitations (vector dbs of the internet solves that)
> Patents and byzantine regulations will stunt its potential
Thats the magic of this technology, its like AWS for highly levered niche intelligence. This arms an entire generation of rebels (entrepreneurs & scientists) to wage a war against big pharma and the FDA.
As an aside, this is why I'm convinced AI & automation will unleash more jobs and productivity like nothing we've seen before. We are at the precipice of a Cambrian explosion! Also why the luddites needs to be shunned.
Re: We are beginning to roll out new voice and image capabilities in ChatGPT
#108Now just throw this into a humanoid looking robot with fine motor skills and we are halfway to a dystopian hellscape that is now only years away instead of decades. What a time to be alive.
Re: We are beginning to roll out new voice and image capabilities in ChatGPT
#109Re: We are beginning to roll out new voice and image capabilities in ChatGPT
#110Earlier quoted context omitted.
I still don't understand how you can talk to something that doesn't provide factual information and just take it at face value? The other day I asked it about the place I live and it made up nonsense, I was trying to get it to help me with an essay and it was just wrong, it was telling me things about this region that weren't real. Do we just drive through a town, ask for a made up history about it and just be satisf…
What LLMs have made me realize more than anything is that we just don't care that much the information we receive being completely factual. I have tried to use it many times to learn a topic, and my experience has been that it is either frustratingly vague or incorrect. It's not a tool that I can completely add to my workflow until it is reliable, but I seem to be the odd one out.
I find this highly concerning but I feel similar.
Even "smart people" I work with seem to have gulped down the LLM cool aid because it's convenient and it's "cool".
Sometimes I honestly think: "just surrender to it all, believe in all the machine tells you unquestionably, forget the fact checking, it feels good to be ignorant... it will be fine...".
I just can't do it though.