Live data from Hacker News

We are beginning to roll out new voice and image capabilities in ChatGPT

openai.com

731–740 of 914 posts

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#731
post #327
post #159

Earlier quoted context omitted.

> The most extreme I can think of is when I want to find when a show comes out and I have to read 10 paragraphs from 5 different sites to realize no one knows. I found that you can be pretty sure no one knows if it’s not already right on the results page. And if the displayed quote for a link on the results page is something like “wondering when show X is coming out?”, then it’s also a safe bet that clicking that lin…

I don't disagree but having to have a learning phase for patterns sounds a bit like people clinging to an old way of things.

You mean like prompt engineering?

What you’re describing as “clinging to an old way of things” is how every single thing has been, ever, new or old.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#732
post #29

This announcement seem to have killed so many startups that were trying to do multi-modal on top of ChatGPT. The way it's progressing with solving use cases with images and voice, not too far when it might be the 'one app to rule them all'. I can already see "Alexa/Siri/Google Home" replacement, "Google Image Search" replacement, ed-tech startups that were solving problems with AI using by taking a photo are also doo…

> This announcement seem to have killed so many startups that were trying to do multi-modal on top of ChatGPT.

Rather than die, why not just pivot to doing multi-modal on top of Llama 2 or some open source model or whatever? It wouldn’t be a huge change

A lot of businesses/governments/etc can’t use OpenAI due to their own policies that prohibit sending their data to third party services. They’ll pay for something they can run on-premise or in their own private cloud

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#733

Earlier quoted context omitted.

Talking to Google and Siri has been positively frustrating this year. On long solo drives, I just want to have a conversation to learn about random things. I've been itching to "talk" to chatGPT and learn more (french | music theory | history | math | whatever) all summer. This should hit the spot!

I've replaced my voice google assistant searches with the voice feature of the Bing app. It's a night and day difference. Bing voice is what I always expected from an AI companion of the future, it is just lacking commands -- setting tasks, home automation, etc.

Did you find a way to do this seamlessly including being able to say something like "Hey Bing", or do you just have a shortcut or widget for this?

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#734
post #61

Earlier quoted context omitted.

Talking to Google and Siri has been positively frustrating this year. On long solo drives, I just want to have a conversation to learn about random things. I've been itching to "talk" to chatGPT and learn more (french | music theory | history | math | whatever) all summer. This should hit the spot!

I still don't understand how you can talk to something that doesn't provide factual information and just take it at face value? The other day I asked it about the place I live and it made up nonsense, I was trying to get it to help me with an essay and it was just wrong, it was telling me things about this region that weren't real. Do we just drive through a town, ask for a made up history about it and just be satisf…

A human driving buddy can make up a lot of stuff too. Have an interesting conversation but don't take it too seriously. If you're really researching something serious then take a mental note to double check things later, pretend as if you're talking to a semi-reliable human who knows a lot but occasionally makes mistakes.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#735

Earlier quoted context omitted.

Then Demis Hassabis ( Deepmind CEO ) probably worked on the tech while he was at LionHead as lead AI programmer on B&W.

Demis was only briefly at LH he went to found Elixir and made Revolution. I believe Richard Evans did the majority of AI in B&W, and he is also at DeepMind now though (assuming it is not just a person with the same name)

> made Revolution

.... which fell far short of his claims, and bombed.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#736

Earlier quoted context omitted.

>And it is not true that computing power will continue to reduce; Moore's Law has been dead for some time now, and if incremental growth in LLMs require exponential growth in computing power the marginal difference won't matter. This is a lot of unfounded assumptions. You don't need Moore's Law. GPU's are not really made with ML training in mind. You don't need exponential growth for anything. The money Open ai spent…

I think you fundamentally don't understand the nature of exponential growth, and the power of diminishing returns. Even if you double the GPU capacity over the next year, you won't even remotely begin to come close enough to producing a step-level growth of capability such as what we experienced between 2 to 3, or even 3 to 4. The LLM concept can only take you so far, and we're approaching the limits of what an LLM i…

>The LLM concept can only take you so far, and we're approaching the limits of what an LLM is capable of.

You don't know that. This is literally just an assertion. An unfounded one at that.

If you couldn't predict how far in 2017 the LLM concept would take us today, then you definitely have no idea how far it could actually go.

>believes we are approaching the limits of LLM size for size’s sake

Nothing to do with thinking they wouldn't improve from scale.

https://web.archive.org/web/20230531203946/https://humanloop...

An interview from Altman later clarifying.

"6. The scaling laws still hold Recently many articles have claimed that “the age of giant AI Models is already over”. This wasn’t an accurate representation of what was meant.

OpenAI’s internal data suggests the scaling laws for model performance continue to hold and making models larger will continue to yield performance. The rate of scaling can’t be maintained because OpenAI had made models millions of times bigger in just a few years and doing that going forward won’t be sustainable. That doesn’t mean that OpenAI won't continue to try to make the models bigger, it just means they will likely double or triple in size each year rather than increasing by many orders of magnitude"

Yes there are economic compute walls. But that's the kind of problem you want, not "innovation".

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#737

Earlier quoted context omitted.

> What LLMs have made me realize more than anything is that we just don't care that much the information we receive being completely factual. I find this highly concerning but I feel similar. Even "smart people" I work with seem to have gulped down the LLM cool aid because it's convenient and it's "cool". Sometimes I honestly think: "just surrender to it all, believe in all the machine tells you unquestionably, forge…

The smart people I've seen using ChatGPT always double check the facts it gives. However, the truth is that RLHF works well to extinguish these lies over time. As more people use the platform and give feedback, the thing gets better. And now, I find it to be pretty darn accurate.

I don't know. The other day I was asking about a biology topic and it straight up gave me a self-contradicting chemical reaction process description. It kept doing that after I pointed out the contradiction. Eventually I got out of this hallucination loop by resetting the conversation and asking again.

It's smart but can also be very dumb.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#738
post #600

Earlier quoted context omitted.

That bike example seemed a mix of underwhelming (for being the demo video) and even confusing. 1. It's not smart enough to recognize from the initial image this is a bolt style seat lock (which a human can). 2. The manual is not shown to the viewer, so I can't infer how the model knows this is a 4mm bolt (or if it is just guessing given that's the most likely one). 3. I don't understand how it can know the toolbox is…

The bike shown in the first image is Specialized Sirrus X. You can make out from the image of the manual that it says "spacer/axle/bolt specifications". Searching for this yields the following Specialized bike manual which is similar: https://www.manualslib.com/manual/1974494/Specialized-Epic-E... -- there are some notable differences, but the Specialized Sirrus X manuals that are online aren't in the same style. The…

Ah good find. yah, I tried bing and it is able to read a photo of that manual page and understand that the seat collar takes a 4mm hex wrench (though hallucinated and told me the torque was 5 Nm, unlike the correct 6.2, suggesting table reading is imperfect).

Toolbox: I just found it too strong to claim you have the right tool, when it really doesn't know that. :)

In the end it does feel like the image reader is just bolted onto an LLM. Basically, just doing object recognition and dumping features into the LLM prompt.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#739

Earlier quoted context omitted.

>And it is not true that computing power will continue to reduce; Moore's Law has been dead for some time now, and if incremental growth in LLMs require exponential growth in computing power the marginal difference won't matter. This is a lot of unfounded assumptions. You don't need Moore's Law. GPU's are not really made with ML training in mind. You don't need exponential growth for anything. The money Open ai spent…

I think you fundamentally don't understand the nature of exponential growth, and the power of diminishing returns. Even if you double the GPU capacity over the next year, you won't even remotely begin to come close enough to producing a step-level growth of capability such as what we experienced between 2 to 3, or even 3 to 4. The LLM concept can only take you so far, and we're approaching the limits of what an LLM i…

Moreover you keep saying we can't scale infinitely. Sure...but nobody is saying we have to. 4 is not as scaled from 3 as 3 was from 2. Doesn't matter, still massive gap.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#740

Earlier quoted context omitted.

Completely agree, latency is key for unlocking great voice experiences. Here's a quick demo I'm working on for voice ordering https://youtu.be/WfvLIEHwiyo Total end-to-end latency is a few hundred milliseconds: starting from speech to text, to the LLM, then to a POS to validate the SKU (no hallucinations are possible!), and finally back to generated speech. The latency is starting to feel really natural. Building out…

That demo is pretty slick. What happens when you go totally off book? Like, ask it to recite the numbers of pi? Or if you become abusive? Will it call the cops?

It's trained to ignore everything else. That way background conversations are ignored as well (like your kids talking in the back of the car while you order).
Post reply on HN