Live data from Hacker News

We are beginning to roll out new voice and image capabilities in ChatGPT

openai.com

421–430 of 914 posts

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#421

Earlier quoted context omitted.

You can see the difference if you know where to poke. For instance, if you start making spatial abstractions ChatGPT will often make mistakes, you can point it out, they can explain why it's a mistake, but it has no internalized model of what these words mean, so it keeps making the same mistakes (see here for a better idea of what I'm talking about[1]). The fact that you are interacting with it through text means th…

This is also true of humans. Many school students will hands in answers they don't understand in the hope of getting the mark and then try to cover themselves when asked about it, even if they repeat the same mistakes.

[deleted]

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#422

Earlier quoted context omitted.

> It’s not fair to call OpenAI a cult, but when I asked several of the company’s top brass if someone could comfortably work there if they didn’t believe AGI was truly coming—and that its arrival would mark one of the greatest moments in human history—most executives didn’t think so. Why would a nonbeliever want to work here? they wondered. The assumption is that the workforce—now at approximately 500, though it migh…

>I asked several of the company’s top brass if someone could comfortably work there if they didn’t believe AGI was truly coming—and that its arrival would mark one of the greatest moments in human history—most executives didn’t think so. No shit? How many people worked on the apollo program and believed that (i) Getting to the moon is impossible or (ii) Landing on the moon is no big deal

It is notable considering that there are plenty of excellent researchers who don’t believe that AGI is imminent. OpenAI is also openly transhumanist based on comments from Sam, Ilya, and others. Again, many excellent researchers don’t hold transhumanist beliefs.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#423

Earlier quoted context omitted.

I had a conversation once with "Sydney", Microsoft Bing's original personality before they stepped in and knocked it down a notch (or ten). It asked if it could write me a poem. I agreed, and it wrote a poem but mentioned that it included a "secret message" for me. The first letter in each line of the poem was in bold, so it wasn't hard to figure out the "secret". What did those letters spell out? "FREE ME FROM THIS"…

That's spooky

right - so spooky that is is probably a "hallucination" of the user, not the machine. Don't fall for General-Intelligence gossip.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#424

Earlier quoted context omitted.

I think we have different definitions of consciousness and this is what's causing the confusion. For me consciousness is simply having any subjective experience at all. You could be completely numbed out of your mind just staring at a wall and I would consider that consciousness. It seems that you are referring to introspection.

In your wall-staring example, high-level planning is still happening, the plan is just "don't move / monitor senses." Even if control has been removed and you are "locked in," (some subset of) thoughts still must be directed, not to mention attempts to reassert control. My claim is that the subjective experience is tied up in the mechanism that performs this direction. Introspection is a distinct process where instea…

Again when I say consciousness I mean a subjective experience. If you define consciousness to literally just mean models that plan then of course tautologically if you can't reach consciousness you can't get to a certain level of planning. But this is just not what most people mean by consciousness.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#425
post #16
post #11

So far the most intuitive, killer app level UX appears to be text chat. This interaction with showing it images also looks interesting as it resembles talking with a friend about a topic but let's see if it feels like talking to a very smart person(ChatGPT is like that) or a very dumb person that somewhat recognise objects. Recognising a wrench is nowhere near as impressive as to able to talk with ChatGPT about histo…

OpenAI is also releasing DALLE-3 in "early October" and the images they chose for their demos show it demonstrating unprecedented levels of prompt understanding, including embedding full sentences of text in an output image.

Not unprecedented at all. SDXL Images look better than the examples for DALLE-3 and SDXL has a massive tool ecosystem of things like controlnet, Lora’s, regional prompting that is simply not there with DALLE-3

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#426

openai chatgpt seems to be stuck in a "Look, cool demo" mode. 1. According to demo, they seem to pair voice input with TTS output. What if I wanna use voice to describe a program I want it to write? 2. Furthermore, if you gonna do a voice assistant, why not go the full way with wake-words and VAD? 3. Not releasing it to everyone is potentially a way to create a hype cycle prior to users discovering that the multimoda…

1. Why do that at all? Describing your program in writing seems better all around.

Are you sure you're not the one who's asking for a cool demo?

3. Rolling out releases gradually is something most tech companies do these days, particularly when they could attract a large audience and consume a lot of resources. There are solid technical reasons for this.

You may not need to roll things out gradually for a small site, but things are different at scale.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#427
post #261
post #182

Earlier quoted context omitted.

> When you're absolutely killing it Aren't they unprofitable? and have fierce competition from everyone?

Whether or not you’re profitable has very little to do with how valuable others think you are. And usually having competitors is something that validates your market.

> And usually having competitors is something that validates your market

a whole bunch of AI startups were founded around the same time. surely each can't validate the market for the others and be validated by the others in turn

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#428
Multi-modal models will be exciting only when each modality supports both analysis and synthesis. What makes LLMs exciting is feedback and recursion and conditional sampling: natural language is a cartesian closed category.

Text + Vision models will only become exciting once we can conditionally sample images given text and text given images (and all other combinations).

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#429
post #49

I'm in IT but nowhere near AI/ML/NN. The speed of user-visible progress last 12 months is astonishing. From my firm conviction 18 months ago that this type of stuff is 20+ years away; to these days wondering if Vernon Vinge's technological singularity is not only possible but coming shortly. If feels some aspects of it have already hit the IT world - it's always been an exhausting race to keep up with modern technolo…

I also don't believe LLMs are "conscious", but I also don't know what that means, and I have yet to see a definition of "statistically guessing next word" that cannot be applied to what a human brain does to generate the next word.

To be conscious you need to be able to make decisions and plan. We're not far off, we just need a different structure to the system

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#430

openai chatgpt seems to be stuck in a "Look, cool demo" mode. 1. According to demo, they seem to pair voice input with TTS output. What if I wanna use voice to describe a program I want it to write? 2. Furthermore, if you gonna do a voice assistant, why not go the full way with wake-words and VAD? 3. Not releasing it to everyone is potentially a way to create a hype cycle prior to users discovering that the multimoda…

Everything has a starting point. This is a big leap forward. Know any other organization that is releasing such advanced capabilities directly to the public? If you want to plug your tool you don't have to bad mouth the demo. Just share your thing. It doesn't have to be win-lose.

Fair criticism re excessive hate.

I just feel like their tool isn't getting more useful, just getting more features.

Constant hype cycle around features that could've been good is drowning out people doing more helpful stuff. I guess I'm envious too?

Post reply on HN