Live data from Hacker News

We are beginning to roll out new voice and image capabilities in ChatGPT

openai.com

591–600 of 914 posts

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#591

Earlier quoted context omitted.

> What LLMs have made me realize more than anything is that we just don't care that much the information we receive being completely factual. I find this highly concerning but I feel similar. Even "smart people" I work with seem to have gulped down the LLM cool aid because it's convenient and it's "cool". Sometimes I honestly think: "just surrender to it all, believe in all the machine tells you unquestionably, forge…

The smart people I've seen using ChatGPT always double check the facts it gives. However, the truth is that RLHF works well to extinguish these lies over time. As more people use the platform and give feedback, the thing gets better. And now, I find it to be pretty darn accurate.

> The smart people I've seen using ChatGPT always double check the facts it gives.

I don't like being told lies in the first place and having to unlearn it.

It doesn't help that I might as well have just gone straight to the "verification" instead.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#592
post #570

Earlier quoted context omitted.

This is not a coincidence, it's increasingly evident that roughly 90% of humans are NPCs.

This is the classic teenage thought of sitting in a bus / subway looking at everyone thinking they're sheep without their own thoughts or much awareness. For everyone who we think is an NPC, there are people who think we are the NPCs. This way of thinking is boring at best, but frankly can be downright dangerous. Everyone has a rich inner world despite shallow immature judgements being made.

Exactly. Most people aren't good at communicating their thoughts or what they see in their mind's eye. These new AI programs will help the average person communicate those, so I'm exciting to see what people come up with. The average person has an amazing mind compared to other animals (as far as we know)

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#593

I went from being worried to thinking it won't replace me anytime soon after using GPT4 for a while and now I'm back to being worried. Because the pace of development is intense. I would love to be financially independent and watch this with excitement and perhaps take on risky and fun projects. Now I'm thinking - how do I double or triple my income so that I reach financial independence in 3 years instead of 10 year…

I'm not convinced that this pace will continue. We're seeing a lot of really cool, rapid evolution of this tech in a short amount of time, but I do think we'll hit a soft ceiling in the not too distant future as well. If you look at something like smartphones, for example. Smartphones, from my perspective, got drastically better and better from about ~2006-2015 or so. They were rapidly improving cameras and battery l…

Maybe that's true but I honestly don't think we can reason at all about how this will progress from a consumer hardware product like the iPhone.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#594
post #563

Earlier quoted context omitted.

I'm very worried constantly. This is the story of the bear, where you just have to be faster than the other guy. For now. The bear is getting faster and faster and it won't be long before it eats all of us. It feels like we're at the end of history. I don't know where we go from here but what are we useful for once this thing is stuck inside a robot like what Tesla is building? What is the point of humanity? Even tak…

re: UBI. I don't think they'll let us starve, but that's a very low bar. If we all become fungible and invaluable they can just feed us Soylent green.

Who is they?

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#595
post #540

The number of comments here of people fearing there is a ghost in the shell is shocking. Are we really this emotional and irrational? Folks, let's all take a moment to remember that AI is nowhere near conscious. It's an illusion based in patterns that mimic humans.

Why is the barrier for so many "consciousness"? Why does it matter whether it's conscious or not if its pragmatic functionality builds use cases that disrupt social contracts (we soon can't trust text, audio OR video - AND we can have human-like text deployed at incredible speed and effectivity), the status quo itself (job displacement), legal statutes and charter (questioning copyright law), and even creativity/self-expression (see: Library of Babel).

When all of this is happening from an unconscious being, why do I care if it's unconscious?

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#596

Earlier quoted context omitted.

I don’t think any of this materially changes job outlook for software development over the next decade. I use ChatGPT daily for school, and used Copilot daily for software development; it gets a lot wrong a lot of the time, and can’t retain necessary context that is critical for being useful long term. I can’t even get it to consume an entire chapter at once to generate notes or flashcards yet. It may slightly change…

This feels fairly naive, ignoring how much progress has happened over the (short) span of one year. This doesn't sound like that tough of a gap to close in another year (again, projecting based off recent progress).

The opposite is fairly naive. Software development is not only dumping tokens into a text file. To have a significant impact on the market, it should do much, much, much more: compile and test code, automatically assess the quality of what its done, be aware of the current design trends (if in UI/UX), ideally innovate, it should also be able to run a debugger, inspect all the variables, and deduce from there how it got something wrong, sometimes with tiny clues that I don't even know how it would get its information (e.g. in graphics programming where you have to actually see at a high frame rate). Oh snap a library is broken ? The AI needs to search online why it's broken, then find a fix (log onto a website to communicate with support, install a missing dep...). It can't be fixed ? Then the AI needs to explain this to the manager, good luck for that. It would need to think and feel like a human, otherwise producing uncanny content that will be either boring, either creepy.

You can think about your daily job and break down all the tasks, and you'll quickly realize that replacing all this is just a monstrous task.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#597
post #9

Earlier quoted context omitted.

Its truly an amazing time to be alive. I'm right there with you, super excited about this decade. Especially what we could do in medicine.

Statistical diagnoses models have offered similar possibilities in medicine for 50 years. Pretty much, the idea is that you can get a far more accurate diagnosis if you take into account the medical history of everyone else in your family, town, workplace, residence and put all of it into a big statistical model, on top of your symptoms and history. However, medical secrecy, processes and laws prevent such things, ev…

Nonsense.

The medical possibilities that will be unlocked by large generative deep multimodal models are on an entirely different scale from "statistical diagnoses." Imagine feeding in an MRI image, asking if this person has cancer, and then asking the model to point out why it thinks the person has cancer. That will be possible within a few years at most. The regulatory challenges will be surmounted eventually once it becomes exceedingly obvious in other countries how impactful this technology is.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#598
post #204

Earlier quoted context omitted.

I think one mean difference in LLM, is what Micheal Scott said in The Office: "Sometimes I'll start a sentence, and I don't even know where it's going. I just hope I find it along the way. Like an improv conversation. An improversation" Human will know what they want to express, choosing words to express it might be similar to LLM process of choosing words, but for LLM it doesn't have that "Here is what i know to exp…

I can only speak from my own internal experience, but don’t your unspoken thoughts take form and exist as language in your mind? If you imagine taking the increasingly common pattern to “think through the problem before giving your answer”, but hiding the pre-answer text from the user, then it seems like that would pretty analogous to how humans think before communicating.

Mine do, but not so much in words. I feel as though my brain has high processing power, but a short context length. When I thought to respond to this comment, I got an inclination something could be added to what I see as an incomplete idea. The idea being humans must form a whole answer in their mind before responding. In my brain it is difficult to keep complex chains juggling around in there. I know because whenever I code without some level of planning it ends up taking 3x longer than it should have.

As a shortcut my brain "feels" something is correct or incorrect, and then logically parse out why I think so. I can only keep so many layers in my head so if I feel nothing is wrong in the first 3 or 4 layers of thought, I usually don't feel the need to discredit the idea. If someone tells me a statement that sounds correct on the surface I am more likely to take it as correct. However, upon digging deeper it may be provably incorrect.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#599

Earlier quoted context omitted.

I believe that the distinguishing factor between what an LLM and a human brain do to generate the next word is that the human brain expresses intentionality originating from inner states and future expectations. As I type this comment I'm sure one could argue that the biological neural networks in my brain are choosing the next word based on statistical guessing, and that the initial prompt was your initial comment.…

> What sets my brain apart from an LLM though is that I am not typing this because you asked me to do it, nor because I needed to reply to the first comment I saw. I am typing this because it is a thought that has been in my mind for a while and I am interested in expressing it to other human brains, motivated by a mix of arrogant belief that it is insightful and a wish to see others either agreeing or providing reas…

The trigger is clearly https://xkcd.com/386/

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#600
post #100

Okay the bike example is cute and impressive, but the human interaction seems to be obfuscating the potentially bigger application. With a few tweaks this is a general purpose solver for robotics planning. There are still a few hard problems between this and a working solution, but it is one of hard problems solved. Will we be seeing general purpose robots performing simple labor powered by chatgpt within the next ha…

That bike example seemed a mix of underwhelming (for being the demo video) and even confusing. 1. It's not smart enough to recognize from the initial image this is a bolt style seat lock (which a human can). 2. The manual is not shown to the viewer, so I can't infer how the model knows this is a 4mm bolt (or if it is just guessing given that's the most likely one). 3. I don't understand how it can know the toolbox is…

The bike shown in the first image is Specialized Sirrus X. You can make out from the image of the manual that it says "spacer/axle/bolt specifications". Searching for this yields the following Specialized bike manual which is similar: https://www.manualslib.com/manual/1974494/Specialized-Epic-E... -- there are some notable differences, but the Specialized Sirrus X manuals that are online aren't in the same style.

The prior page (8) shows "SEAT COLLAR 4mm HEX" and, based on looking up seat collar in an image search, the part in question matches.

In terms of the toolbox, note that it only identified the location of the Allen wrench set. The advice was just "Within that set, find the 4 mm Allen (Hex) key". Had they replied with "I don't see any sizes in mm", the conversation could've continued with "Your Allen keys might be using SAE sizing. A compatible size will be 5/32, do you see that in your set?"

Post reply on HN