Live data from Hacker News

We are beginning to roll out new voice and image capabilities in ChatGPT

openai.com

701–710 of 914 posts

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#701
post #100

Okay the bike example is cute and impressive, but the human interaction seems to be obfuscating the potentially bigger application. With a few tweaks this is a general purpose solver for robotics planning. There are still a few hard problems between this and a working solution, but it is one of hard problems solved. Will we be seeing general purpose robots performing simple labor powered by chatgpt within the next ha…

That bike example seemed a mix of underwhelming (for being the demo video) and even confusing. 1. It's not smart enough to recognize from the initial image this is a bolt style seat lock (which a human can). 2. The manual is not shown to the viewer, so I can't infer how the model knows this is a 4mm bolt (or if it is just guessing given that's the most likely one). 3. I don't understand how it can know the toolbox is…

Right. It appeared that the response to the first image and question would have been the same if the image wasn't provided.

I wasn't impressed with the demo but we'll see what real world results get.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#702
post #157

Earlier quoted context omitted.

You listen to Joe Rogan with the idea that this is a normal dude talking not an expert beyond martial arts and comedy. A person who uses ChatGPT must have the understanding that it's not like Google search. The layman, however, has no idea that ChatGPT can give coherent incorrect information and treats the information as true. Most people won't use it for infotainment and OpenAI will try its best to downplay the hall…

Give people more credit. If you're using an AI these days, you have to know it hallucinates sometimes. There's even a warning about it when you log in.

There's a contingent of the population passing videos around on tiktok genuinely concerned that AIs have a mind of their own

no I will not give the public credit, most people have no grounding to discern wtf a language model is and what it's doing, all they know is computers didn't use to talk and now they do

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#703

Voice has the potential to be awesome. This demo is really underwhelming to me because of the multi-second latency between the query and response, just like every other lame voice assistant. It doesn't have to be this way! I have a local demo using Llama 2 that responds in about half a second and it feels like talking to an actual person instead of like Siri or something. I really should package it up so people can t…

Completely agree, latency is key for unlocking great voice experiences. Here's a quick demo I'm working on for voice ordering https://youtu.be/WfvLIEHwiyo Total end-to-end latency is a few hundred milliseconds: starting from speech to text, to the LLM, then to a POS to validate the SKU (no hallucinations are possible!), and finally back to generated speech. The latency is starting to feel really natural. Building out…

The voice does not seem to be able to pronounce the L in “else”. What’s happening there?

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#704
post #521

Earlier quoted context omitted.

The real problem is distribution of the output of production. We will need something like UBI eventually.

UBI is just not happening any time soon in the US. To start, half of the country is already default against it. Precisely 0 people in Congress, the White House, or those in adjacent power roles (lobbyists and whatnot) are for it or have any idea what it is. Aside from rolling out the guillotine, I don't see UBI a possibility until the 2nd half of the 21st century. There's just too many forces and entities alive that…

I think the plan is first robots take our jobs, then UBI. If you gave people free money now we'd be suffering from a lack of workers due to general robot non existence. I'm guessing 2045 maybe?

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#705

Earlier quoted context omitted.

UBI is a bandaid on top of capitalism. It is saying "we have a system where people die if they don't have money, so we'll give people money." It's not a real fix. A real fix would be replacing the system with one where people don't need money in order to not die. We're going to keep automating more and more things. I think that much is inevitable. Eventually, we may get to a point where very few jobs are necessary fo…

That’s not going to come to fruition, and no amount of dreamy socialist fanfiction’ing is going to make it so. People pay for value. Produce value for others, get paid. LLM’s are tools to make humans able to produce more value, and will not replace humans, although the job market will change, and hopefully utilize humans better. People, NOT machines, are the ultimate judgers of what is valuable and the ultimate produ…

> Produce value for others, get paid

So if a human is unable to produce value, they don't get (food/education/heathcare/)? That seems to be the implication. We in developed countries already have some amount of "value risk hedging" (I'm loathe to say "socialism" here), we just disagree endlessly how much is the optimal amount. But we've determined that wards of the state, universal education, and some amount of food support for the poor is the absolute bare minimum for a developed society.

> People, NOT machines, are the ultimate judgers of what is valuable and the ultimate producers of value.

Uhhh we already have software which sifts through resumes to allow/reject candidates, before it gets to any kind of human judge, so we are already gating value assessments.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#706
post #83

Earlier quoted context omitted.

Joe Rogan has made tons of money off talking without providing factual information. Hollywood has also made tons of money off movies "inspired by real events" that hallucinate key facts relevant to the movie's plot and characters. There's a huge market for infotainment that is "inspired by facts" but doesn't even try to be accurate.

OpenAI isn't marketing ChatGPT as, "infotainment."

now that you mention it, a big "for entertainment purposes only" banner like they use to have on all the psychic commercials on tv would not be inappropriate. it's incredible that LLMs are being marketed as general purpose assistants with a tiny asterisk, "may contain inaccuracies" like it's a walnut contamination

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#707

I went from being worried to thinking it won't replace me anytime soon after using GPT4 for a while and now I'm back to being worried. Because the pace of development is intense. I would love to be financially independent and watch this with excitement and perhaps take on risky and fun projects. Now I'm thinking - how do I double or triple my income so that I reach financial independence in 3 years instead of 10 year…

[deleted]

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#709
post #83

Earlier quoted context omitted.

Joe Rogan has made tons of money off talking without providing factual information. Hollywood has also made tons of money off movies "inspired by real events" that hallucinate key facts relevant to the movie's plot and characters. There's a huge market for infotainment that is "inspired by facts" but doesn't even try to be accurate.

Wait until you learn about the mainstream media.

Rogan is literally the largest podcast on the Spotify. It's the definition of mainstream.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#710

Earlier quoted context omitted.

>the gap between 3 and 4 is not anything as large as the gap between 2 and 3. We'll just have to agree to disagree. 3 was a signal of things to come but it was ultimately a bit of a toy, a research curiosity. Utility wise, they are worlds apart. >if there really was no innovation at all as you claim, then there will be a limit hit at computing capability and cost. computing capability and cost are just about the one…

Do you realize I'm not disagreeing with you about the difference between 3 and 4? Reread what I wrote. I contrasted 3 and 4 with 2 and 3, which you seem to be entirely ignoring. 3 and 4 could be worlds apart, but wouldn't matter if 2 and 3 were two worlds apart, for example. And it is not true that computing power will continue to reduce; Moore's Law has been dead for some time now, and if incremental growth in LLMs…

>And it is not true that computing power will continue to reduce; Moore's Law has been dead for some time now, and if incremental growth in LLMs require exponential growth in computing power the marginal difference won't matter.

This is a lot of unfounded assumptions.

You don't need Moore's Law. GPU's are not really made with ML training in mind. You don't need exponential growth for anything. The money Open ai spent on GPT-4 a year ago could train a model twice as large today. and that amount is a drop in the bucket for the R&D of large corporations. Microsoft gave open ai 10B. amazon gave anthropic 4B

>So compute will not fall at the rate you would need it to for LLMs to actually compete in any meaningful way with human software engineers.

I don't think the compute reuired is anywhere near as much as you think it is.

https://arxiv.org/abs/2309.12499

>We are not guaranteed to continue to progress in anything just because we have in the past.

Nothing is guaranteed. But the scaling plots show no indication of a slow down so it's up to you to provide a concrete reason this object in motion is going to stop immediately and conveniently right now. If all you have is "well it just can't keep getting better right" then visit the 2 and 3 threads to see how meaningless such unfounded assertions are.

Post reply on HN