Live data from Hacker News

We are beginning to roll out new voice and image capabilities in ChatGPT

openai.com

721–730 of 914 posts

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#721

Voice has the potential to be awesome. This demo is really underwhelming to me because of the multi-second latency between the query and response, just like every other lame voice assistant. It doesn't have to be this way! I have a local demo using Llama 2 that responds in about half a second and it feels like talking to an actual person instead of like Siri or something. I really should package it up so people can t…

Completely agree, latency is key for unlocking great voice experiences. Here's a quick demo I'm working on for voice ordering https://youtu.be/WfvLIEHwiyo Total end-to-end latency is a few hundred milliseconds: starting from speech to text, to the LLM, then to a POS to validate the SKU (no hallucinations are possible!), and finally back to generated speech. The latency is starting to feel really natural. Building out…

That demo is pretty slick. What happens when you go totally off book? Like, ask it to recite the numbers of pi? Or if you become abusive? Will it call the cops?

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#722

Earlier quoted context omitted.

There is clearly a plateau in how good a ux can be. It might be a local optimum but still you solve the task the user wants. I don't see a clear ceiling in intelligence. And if the ceiling is how much of the human tasks can be replaced then I think when we reach it the world is going to look very different from now. (Let's also not discount how much the world changed since the introduction of the smartphone.)

> I don't see a clear ceiling in intelligence The plateau in this case is presumably how far you can advance intelligence from the current model architectures. There seems to be diminishing returns from throwing more layers, parameters or training data at these things. We will see improvements but for dramatic increases I think we'll need new breakthroughs. New inventions are hard to predict, pretty much by definitio…

That's more or less what I was getting at; the cool new GANN and LLM models have a certain set of problems that they will solve exceptionally well, and then another set of problems that they will solve "pretty well", but I don't think they'll solve every problem.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#723

Earlier quoted context omitted.

Voice assistants have always been a half complete product. They were shown off as a cool feature, then they were never integrated so they were useful. The two biggest features I want are for the voice assistants to read something for me, and to do something on google/Apple Maps hand free. Neither of these ever work. “Siri/ ok google add the next gas station on the route” or “take me to the Chinese restaurant in Hobok…

In the current world: Me: “OK Google, take me to the Chinese restaurant in Hoboken” Google Assistant: “Calling Jessica Hobkin”.

This reminds me of ordering at a drive through with a human at times:

"I'd like an iced tea" "An icee?" "No an iced tea" "Hi-C?"

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#724

Earlier quoted context omitted.

Last I heard, OpenAI was losing massive amounts of money to run all this. Has that changed? Because past history shows that the first out of the gate is not the definitive winner much of the time. We aren't still using gopher. We aren't searching with altavista. We don't connect to the internet with AOL. AI is going to change many things. That is all the more reason to keep working on how best to make it work, not gi…

you're absolutely right. also, I did not know until today's thread that OpenAI's stated goal is building AGI. which is probably never going to happen, ever, no matter how good technology gets. which means yes, we are absolutely looking at AltaVista here, not Google, because if you subtract a cult from an innovative business, you might be able to produce a profitable business.

Why isn’t AGI ever going to happen? Ever?

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#725

Earlier quoted context omitted.

Do you realize I'm not disagreeing with you about the difference between 3 and 4? Reread what I wrote. I contrasted 3 and 4 with 2 and 3, which you seem to be entirely ignoring. 3 and 4 could be worlds apart, but wouldn't matter if 2 and 3 were two worlds apart, for example. And it is not true that computing power will continue to reduce; Moore's Law has been dead for some time now, and if incremental growth in LLMs…

>And it is not true that computing power will continue to reduce; Moore's Law has been dead for some time now, and if incremental growth in LLMs require exponential growth in computing power the marginal difference won't matter. This is a lot of unfounded assumptions. You don't need Moore's Law. GPU's are not really made with ML training in mind. You don't need exponential growth for anything. The money Open ai spent…

I think you fundamentally don't understand the nature of exponential growth, and the power of diminishing returns. Even if you double the GPU capacity over the next year, you won't even remotely begin to come close enough to producing a step-level growth of capability such as what we experienced between 2 to 3, or even 3 to 4. The LLM concept can only take you so far, and we're approaching the limits of what an LLM is capable of. You generally can't just push an innovation infinitely, it will have a drop-off point somewhere.

the "Large" part of LLMs is probably done. We've gotten as far as we can with those style of models, and the next innovation will be in smaller, more targeted models.

> As costs have skyrocketed while benefits have leveled off, the economics of scale have turned against ever-larger models. Progress will instead come from improving model architectures, enhancing data efficiency, and advancing algorithmic techniques beyond copy-paste scale. The era of unlimited data, computing and model size that remade AI over the past decade is finally drawing to a close. [0]

> Altman, who was interviewed over Zoom at the Imagination in Action event at MIT yesterday, believes we are approaching the limits of LLM size for size’s sake. “I think we’re at the end of the era where it’s gonna be these giant models, and we’ll make them better in other ways,” Altman said. [1]

[0] https://venturebeat.com/ai/openai-chief-says-age-of-giant-ai...

[1] https://techcrunch.com/2023/04/14/sam-altman-size-of-llms-wo...

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#726
post #395
post #180

Earlier quoted context omitted.

Google demoed this a few months ago https://www.deepmind.com/blog/rt-2-new-model-translates-visi...

They are really good at keeping demos as demos

The implementation that manifests itself as an extremely creepy, downright concerning level of dubious moral transgressions isn't nearly as publicly glamorous as their tech demos.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#727
post #61

Earlier quoted context omitted.

I still don't understand how you can talk to something that doesn't provide factual information and just take it at face value? The other day I asked it about the place I live and it made up nonsense, I was trying to get it to help me with an essay and it was just wrong, it was telling me things about this region that weren't real. Do we just drive through a town, ask for a made up history about it and just be satisf…

This is a fairly perpetual discussion, but I'll go for another round: I feel like using LLM today is like using search 15 years ago - you get a feel for getting results you want. I'd never use chatGPT for anything that's even remotely obscure, controversial, or niche. But through all my double-checking, I've had phenomenal success rate in getting useful, readable, valid responses to well-covered / documented topics s…

> I feel like using LLM today is like using search 15 years ago - you get a feel for getting results you want.

I don't think it's quite the same.

With search results, aka web sites, you can compare between them and get a "majority opinion" if you have doubts - it doesn't guarantee correctness but it does improve the odds.

Some sites are also more reputable and reliable than others - e.g. if the information is from Reuters, a university's courseware, official government agencies, ... etc. it's probably correct.

With LLMs you get one answer and that's it - although some like Bard provide alternate drafts but they are all from the same source and can all be hallucinations ...

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#728

Earlier quoted context omitted.

I'm not convinced that this pace will continue. We're seeing a lot of really cool, rapid evolution of this tech in a short amount of time, but I do think we'll hit a soft ceiling in the not too distant future as well. If you look at something like smartphones, for example. Smartphones, from my perspective, got drastically better and better from about ~2006-2015 or so. They were rapidly improving cameras and battery l…

> I do think we'll hit a soft ceiling in the not too distant future ... it's going to plateau and progress will become substantially more gradual. I don't think this will age well. It's a matter of simple compute power to advance from realistic text/token prediction, to realistic synthesis of stuff like human (or animal) body movement, for all kinds of situations, including realistic facial/body language, moods, and…

Isn't video prediction a substantially harder problem than text prediction? At least that was the case a couple of years ago with RNNs/LSTMs. Haven't kept up with the research, maybe there's been progress.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#729

Earlier quoted context omitted.

I disagree that at current possibility it was "totally possible" but it was 100% obvious by that point that it was going to be possible very soon . IMO that has been clear since ~2019.

GPT3 existed. OCR existed. Object recognition existed.

GPT3 was not as good as 3.5. Multimodal is not the same as OCR + object recognition.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#730
post #248

Earlier quoted context omitted.

Talking to Google and Siri has been positively frustrating this year. On long solo drives, I just want to have a conversation to learn about random things. I've been itching to "talk" to chatGPT and learn more (french | music theory | history | math | whatever) all summer. This should hit the spot!

Agreed. After using ChatGPT at all Siri is absolutely frustrating. Example from a couple days ago: Me, in the shower so not able to type: "Hey Siri, add 1.5 inch brad nails to my latest shopping list note." Siri: "Sorry, I can't help with that." ... Really, Siri? You can't do something as simple as add a line to a note in the first-party Apple Notes app?

That’s extra frustrating because Siri absolutely had that functionality at some point in the past, and may even still have it if you say the right incantation. Those incantations change in unpredictable and unknowable ways though.
Post reply on HN