Live data from Hacker News

GPT-4o

openai.com

551–560 of 1001 posts

Re: GPT-4o

#552
post #43
post #24

The most impressive part is that the voice uses the right feelings and tonal language during the presentation. I'm not sure how much of that was that they had tested this over and over, but it is really hard to get that right so if they didn't fake it in some way I'd say that is revolutionary.

(I work at OpenAI.) It's really how it works.

I couldn’t quite tell from the announcement, but is there still a separate TTS step, where GPT is generating tones/pitches that are to be used, or is it completely end to end where GPT is generating the output sounds directly?

Re: GPT-4o

#553
In the customer support example, he tells it his new phone doesn't work, and then it just starts making stuff up like how the phone was delivered 2 days ago, and there's physically nothing wrong with it, which it doesn't actually know. It's a very impressive tech demo, but it is a bit like they are pretending we have AGI when we really don't yet.

(Also, they managed to make it sound exactly like an insincere, rambling morning talk show host - I assume this is a solvable problem though.)

Re: GPT-4o

#554
post #551

Seems that no client-side changes needed for gpt-4o chat completion Added a custom OpenAI endpoint to https://recurse.chat (i built it) and it just works: https://twitter.com/recursechat/status/1790074433610137995

but does it do the full multimodal in-out capability shown in the app :)

Re: GPT-4o

#555

I worry that this tech will amplify the cultural values we have of "good" and "bad" emotions way more than the default restrictions that social media platforms put on the emoji reactions (e.g., can't be angry on LinkedIn). I worry that the AI will not express anger, not express sadness, not express frustration, not express uncertainty, and many other emotions that the culture of the fine-tuners might believe are "bad…

Imagine how warped your personality might become if you use this as an entire substitute for human interaction. Should people use this as bf/gf material we might just be further contributing to decreasing the fertility rate. However we might offset this by reducing the suicide rate somewhat too.

I've worked in emotional communication and conflict resolution for over 10 years and I'm honestly just feeling a huge swirl of uncertainty on how this—LLMs in general, but especially the genAI voices, videos, and even robots—will impact how we communicate with each other and how we bond with each other. Does bonding with an AI help us bond more with other humans? Will it help us introspect more and dig deeper into our common humanity? Will we learn how to resolve conflict better? Will we learn more passive aggression? Become more or less suicidal? More or less loving?

I just, yeah, feel a lot of fear of even thinking about it.

Re: GPT-4o

#556
post #115

The usual critics will quickly point out that LLMs like GPT-4o still have a lot of failure modes and suffer from issues that remain unresolved. They will point out that we're reaping diminishing returns from Transformers. They will question the absence of a "GPT-5" model. And so on -- blah, blah, blah, stochastic parrots, blah, blah, blah. Ignore the critics. Watch the demos. Play with it. This stuff feels magical .…

Funnily, I’d prefer HAL’s unemotional monotone over GPT’s woke hyperbola any second.

Re: GPT-4o

#557
post #115

The usual critics will quickly point out that LLMs like GPT-4o still have a lot of failure modes and suffer from issues that remain unresolved. They will point out that we're reaping diminishing returns from Transformers. They will question the absence of a "GPT-5" model. And so on -- blah, blah, blah, stochastic parrots, blah, blah, blah. Ignore the critics. Watch the demos. Play with it. This stuff feels magical .…

I mean, humans also have tons of failures modes, but we've learned to live them over time.

The average human have tons of quirks, talk over each other all the time, generally can't solve complex problems in a casual conversion setting, and are not always cheery and ready to please like Scarlet's character in Her.

I think our expectations of AI is way too high from our exposure to science fiction.

Re: GPT-4o

#558

In the customer support example, he tells it his new phone doesn't work, and then it just starts making stuff up like how the phone was delivered 2 days ago, and there's physically nothing wrong with it, which it doesn't actually know. It's a very impressive tech demo, but it is a bit like they are pretending we have AGI when we really don't yet. (Also, they managed to make it sound exactly like an insincere, ramblin…

It’s possible to imagine using ChatGPT’s memory, or even just giving the context in an initial brain dump that would allow for this type of call. So don’t feel like it’s too far off.

Re: GPT-4o

#559
post #24

The most impressive part is that the voice uses the right feelings and tonal language during the presentation. I'm not sure how much of that was that they had tested this over and over, but it is really hard to get that right so if they didn't fake it in some way I'd say that is revolutionary.

>The most impressive part is that the voice uses the right feelings and tonal language during the presentation. Consequences of audio2audio (rather than audio >text text>audio). Being able to manipulate speech nearly as well as it manipulates text is something else. This will be a revelation for language learning amongst other things. And you can interrupt it freely now!

I asked it to make a bird noise, instead it told me what a bird sounds like with words. True audio to audio should be able to be any noise, a trombone, traffic, a crashing sea, anything. Maybe there is a better prompt there but it did not seem like it.

Re: GPT-4o

#560
post #24

The most impressive part is that the voice uses the right feelings and tonal language during the presentation. I'm not sure how much of that was that they had tested this over and over, but it is really hard to get that right so if they didn't fake it in some way I'd say that is revolutionary.

"Right" feelings and tonal language? "Right" for what? For whom?

We've already seen how much damage dishonest actors can do by manipulating our text communications with words they don't mean, plans they don't intend to follow through on, and feelings they don't experience. The social media disinfo age has been bad enough.

Are you sure you want a machine which is able to manipulate our emotions on an even more granular and targetted level?

LLMs are still machines, designed and deployed by humans to perform a task. What will we miss if we anthropomorphize the product itself?

Post reply on HN