Live data from Hacker News

We are beginning to roll out new voice and image capabilities in ChatGPT

openai.com

681–690 of 914 posts

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#681

Earlier quoted context omitted.

>What actually was the innovation in LLMs that produced the kind of AI we're seeing now? Is that innovation ongoing or did it happen, and now we're seeing the various optimizations of that innovation? Past the introduction of the transformer in 2017, There is no big "innovation". It is just scale. Bigger models are better. The last 4 years can be summed up that simply. >Is voice and image integration with ChatGPT a w…

Firstly no, the gap between 3 and 4 is not anything as large as the gap between 2 and 3. Secondly, nothing you said here changed as of this announcement. Nothing here makes it any more or less likely LLMs will risk software engineering jobs. Thirdly, you can take what Sam Altman says with as many grains of salt as you like, if there really was no innovation at all as you claim, then there will be a limit hit at compu…

>the gap between 3 and 4 is not anything as large as the gap between 2 and 3.

We'll just have to agree to disagree. 3 was a signal of things to come but it was ultimately a bit of a toy, a research curiosity. Utility wise, they are worlds apart.

>if there really was no innovation at all as you claim, then there will be a limit hit at computing capability and cost.

computing capability and cost are just about the one thing you can bank on to reduce. already training gpt-4 today would be a fraction of the cost than it was when open ai did it and that was just over a year ago.

Today's GPU's take ML into account to some degree but they are nowhere near as calibrated for it as they could be. That work has just begun to start.

Of any of the possible barriers, compute is exactly the kind you want. It will fall.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#682
post #665

Earlier quoted context omitted.

If I took 2 weeks off from work I could build this prototype quite easily. We're in an interesting period where the space of possibilities is so large it just takes a while for the "market" to exhaust it.

Quizlet has a feature to build flashcards using AI. I'm sure they could write a backend service that just chunked the entire chapter.

It doesn't work well enough yet. The flashcards it generates don't actually fit well into its own ecosystem. When you try to build the "quizzes", the wrong answers are trivially spottable. Further, even the generated questions are stilted don't hit parity with manually generated flashcards.

My use of ChatGPT for this purpose is so far mostly limited to a sanity check, e.g. "Do these notes cover the major points of this topic?" Usually it'll spit back out "Yep looks good" or some major missed point, like The Pacific Railway Act of 1862 for a topic on the Civil War's economic complexity.

I'll also use it to reformat content, "Convert these questions and answers into Anki format."

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#683

Earlier quoted context omitted.

> my skillset is being rapidly replaced. Why do you have only one? Learn some trades. AI isn't going to be demolishing a bathroom and installing tile any time soon.

I don't know what your salary is but mine isn't going to be replaced by demoing a bathroom and I have a mortgage and a standard of living I was hoping to be able to afford at least until my kids are out of the house.

Unless you're making a ridiculous amount of money, you can definitely match a developer salary remodeling homes. So long as you're the actual business owner. This was just an example, of course.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#684

I went from being worried to thinking it won't replace me anytime soon after using GPT4 for a while and now I'm back to being worried. Because the pace of development is intense. I would love to be financially independent and watch this with excitement and perhaps take on risky and fun projects. Now I'm thinking - how do I double or triple my income so that I reach financial independence in 3 years instead of 10 year…

I'm not convinced that this pace will continue. We're seeing a lot of really cool, rapid evolution of this tech in a short amount of time, but I do think we'll hit a soft ceiling in the not too distant future as well. If you look at something like smartphones, for example. Smartphones, from my perspective, got drastically better and better from about ~2006-2015 or so. They were rapidly improving cameras and battery l…

Or an exponential perhaps. Like the Wait But Why thing (https://waitbutwhy.com/2015/01/artificial-intelligence-revol... bottom of the article)

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#685

Earlier quoted context omitted.

that's completely apples to oranges. OpenAI is in the business of leveraging the utility of large language models. that's their moon. if they think instead that they're in the business of creating some kind of ridiculous robot god, that is definitely interesting information about them. because that's no moon.

>OpenAI is in the business of leveraging the utility of large language models. No Open AI is in the business of creating their vision of Artificial General Intelligence (which they define as that is generally smarter than humans ) and they believe LLMs are a viable path. This has always been the case. It's not some big secret and they have many posts which talk upon their expectations and goals in this space. https:/…

> No Open AI is in the business of creating their vision of Artificial General Intelligence

that's a project, not a business.

> GPT as a product comes second and it shows

we can agree on that, at least.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#686
post #521

I went from being worried to thinking it won't replace me anytime soon after using GPT4 for a while and now I'm back to being worried. Because the pace of development is intense. I would love to be financially independent and watch this with excitement and perhaps take on risky and fun projects. Now I'm thinking - how do I double or triple my income so that I reach financial independence in 3 years instead of 10 year…

The real problem is distribution of the output of production. We will need something like UBI eventually.

UBI is just not happening any time soon in the US. To start, half of the country is already default against it. Precisely 0 people in Congress, the White House, or those in adjacent power roles (lobbyists and whatnot) are for it or have any idea what it is.

Aside from rolling out the guillotine, I don't see UBI a possibility until the 2nd half of the 21st century. There's just too many forces and entities alive that don't want it

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#687

openai chatgpt seems to be stuck in a "Look, cool demo" mode. 1. According to demo, they seem to pair voice input with TTS output. What if I wanna use voice to describe a program I want it to write? 2. Furthermore, if you gonna do a voice assistant, why not go the full way with wake-words and VAD? 3. Not releasing it to everyone is potentially a way to create a hype cycle prior to users discovering that the multimoda…

[deleted]

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#688

Voice has the potential to be awesome. This demo is really underwhelming to me because of the multi-second latency between the query and response, just like every other lame voice assistant. It doesn't have to be this way! I have a local demo using Llama 2 that responds in about half a second and it feels like talking to an actual person instead of like Siri or something. I really should package it up so people can t…

all it has to do is add a random selection of "uhms" and "ahhs" and "mmm"

Unfortunately, Bark is probably way too slow to use for the TTS portion given the latency concerns or that would be covered.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#689
post #248

Earlier quoted context omitted.

Talking to Google and Siri has been positively frustrating this year. On long solo drives, I just want to have a conversation to learn about random things. I've been itching to "talk" to chatGPT and learn more (french | music theory | history | math | whatever) all summer. This should hit the spot!

Agreed. After using ChatGPT at all Siri is absolutely frustrating. Example from a couple days ago: Me, in the shower so not able to type: "Hey Siri, add 1.5 inch brad nails to my latest shopping list note." Siri: "Sorry, I can't help with that." ... Really, Siri? You can't do something as simple as add a line to a note in the first-party Apple Notes app?

appending to a text file, what do you think this is - unix?

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#690

I know there are shades of grey to how they operate, but the near constant stream of stuff they're shipping keeps me excited. The LLM boom of the last year (Open AI, llama, et al) has me giddy as a software person. It's a reach, but I truly feel like I'm watching the pyramids of our time get made.

Yep. Several months ago I was imagining this exact feature, and yet as I watched a video of it in use, I'm still in awe. It's incredible. I think this could bring back Google Glass, actually. Imagine wearing them while cooking, and having ChatGPT give you active recipe instructions as well as real-time feedback. I could see that within the next 1-3 years.

Related, the iOS app has supported realtime conversations for months now, using Shortcuts app and the "Hey Siri " trigger to initiate it. Mine is "Hey Siri, let's talk".

I think they're using Siri for dictation, though. Using Whisper, especially if they use speaker identification, is going to be great. But, a shortcut will still be required to get it going.

Post reply on HN