Live data from Hacker News

GPT-4o

openai.com

651–660 of 1001 posts

Re: GPT-4o

#651
post #625

I worry that this tech will amplify the cultural values we have of "good" and "bad" emotions way more than the default restrictions that social media platforms put on the emoji reactions (e.g., can't be angry on LinkedIn). I worry that the AI will not express anger, not express sadness, not express frustration, not express uncertainty, and many other emotions that the culture of the fine-tuners might believe are "bad…

Sam Altman talked a little bit about this in his recent appearance on the All-In podcast [0]. I'm paraphrasing, but his vision is that ai assistants in the near term will be like a senior level employee - they'll push back when it makes sense to and not just be sycophants. [0]: https://youtube.com/watch?v=nSM0xd8xHUM

I really wonder how that'll go, because workplaces already seem to limit human communication and emotion to "professional behavior." I'm glad he's thinking about it and I hope they're able to figure out how to improve human communication so that we can resolve conflict with bots. In his example (around 21:05), he talks about how the bot could do something if the person wants but there might be consequences to that action, and I think that makes more sense if the bot is acting like a computer that has limits on what it can do. For example, if I ask it to do two tasks that really stretch its computational limits, I'd hope it would let me know. But if it pretends it's a human with human limits, I don't know how much that'd help, unless it were a training exercise.

Re: GPT-4o

#652
post #569
post #495

We've had voice input and voice output with computers for a long time, but it's never felt like spoken conversation. At best it's a series of separate voice notes. It feels more like texting than talking. These demos show people talking to artificial intelligence. This is new. Humans are more partial to talking than writing. When people talk to each other (in person or over low-latency audio) there's a rich metadata…

But in this case you're not talking with a real person. Instinctively, I dislike a robot that pretends to be a real human being.

To your point, there's been a lot of talk about AI, regulation, guardrails, whatever. Now is the time to say, AI must speak such that we know it's AI and not a real human voice.

We get the upside of conversation, and avoid the downside of falling asleep at the wheel (as Ethan Mollick mentions in "Co-Intelligence".)

Re: GPT-4o

#653

I found these videos quite hard to watch. There is a level of cringe that I found a bit unpleasant. It’s like some kind of uncanny valley of human interaction that I don’t get on nearly the same level with the text version.

The what I presume default female voice and tone speaks like an AI character straight out of a Black Mirror episode.

Re: GPT-4o

#654

I worry that this tech will amplify the cultural values we have of "good" and "bad" emotions way more than the default restrictions that social media platforms put on the emoji reactions (e.g., can't be angry on LinkedIn). I worry that the AI will not express anger, not express sadness, not express frustration, not express uncertainty, and many other emotions that the culture of the fine-tuners might believe are "bad…

> Almost like it might become an AI "yes man." Seems like that ship sailed a long time ago. For social media at least, where for example FB will generally do its best to show you posts that you already agree with. Reinforcing your existing biases may not be the goal but it's certainly an effect.

I appreciate you pointing this out. I think the effect may be even larger when it's not an ad I'm trying to ignore or even a post that was fed to me, but words and emotions that were created specifically for me. Social media seems to find already written posts/images/videos that I may want and put them in front of my face. This would be writing those things directly for me.

Re: GPT-4o

#655
post #591

I worry that this tech will amplify the cultural values we have of "good" and "bad" emotions way more than the default restrictions that social media platforms put on the emoji reactions (e.g., can't be angry on LinkedIn). I worry that the AI will not express anger, not express sadness, not express frustration, not express uncertainty, and many other emotions that the culture of the fine-tuners might believe are "bad…

Try getting GPT to draw pictures of Mohammed and it gets pretty scared.

> Try getting GPT to draw pictures of Mohammed and it gets pretty scared.

Yet, it has no issue drawing cartoons of Jesus. Why the double standard?

Re: GPT-4o

#656
I've worked quite a bit with STT and TTS over the past ~7 years, and this is the most impressive and even startling demo I've seen.

But I would like to see how this is integrated into applications by third party developers where the AI is doing a specific job. Is it still as impressive?

The biggest challenge I've had with building any autonomous "agents" with generic LLM's is they are overly gullible and accommodating, requiring the need to revert back to legacy chatbot logic trees etc. to stay on task and perform a job. Also STT is rife with speaker interjections, leading to significant user frustrations and they just want to talk to a person. Hard to see if this is really solved yet.

Re: GPT-4o

#657

I found these videos quite hard to watch. There is a level of cringe that I found a bit unpleasant. It’s like some kind of uncanny valley of human interaction that I don’t get on nearly the same level with the text version.

All these demo style ads/videos are super jarring and uncanny valley-esque to watch as an Australian. The US corporate cultural norms are super bizarre to the rest of the world, and the California based holy omega of tech companies really takes this to the extreme. The application might work well if you interact with it like you are a normal human being - but I can't tell because this presentation is corporate robots…

Agreed. Americans why are you like this?

Re: GPT-4o

#658

Impressed by the model so far. As far as independent testing goes, it is topping our leaderboard for chess puzzle solving by a wide margin now: https://github.com/kagisearch/llm-chess-puzzles?tab=readme-o...

I wasn't impressed in the first 5 minutes of using it but it is quite impressive after 2 solid hours of random topics.

Much faster for sure but I have also not had anything give an error in python with jupyter. Usually you could only stray so far with more obscure python libraries before it starts producing errors.

That much better than 4 in chess is pretty shocking in a great way.

Re: GPT-4o

#659

I found these videos quite hard to watch. There is a level of cringe that I found a bit unpleasant. It’s like some kind of uncanny valley of human interaction that I don’t get on nearly the same level with the text version.

All these demo style ads/videos are super jarring and uncanny valley-esque to watch as an Australian. The US corporate cultural norms are super bizarre to the rest of the world, and the California based holy omega of tech companies really takes this to the extreme. The application might work well if you interact with it like you are a normal human being - but I can't tell because this presentation is corporate robots…

That was my reaction (as an Australian) too. The AI is so verbose and chirpy by default. There was even a bit in one video where he started talking over the top of the AI because it was rabbiting on.

But I find the text version similar. Delivers too much and too slowly. Just get me the key info!

Re: GPT-4o

#660
> We recognize that GPT-4o’s audio modalities present a variety of novel risks. Today we are publicly releasing text and image inputs and text outputs. Over the upcoming weeks and months, we’ll be working on the technical infrastructure, usability via post-training, and safety necessary to release the other modalities.

So they're using the same GPT4 model with a relatively small improvement, and no voice whatsoever outside of the prerecorded demos. This is not a "launch" or even an announcement. This is a demo of something which may or may not work in the future.

Post reply on HN