Live data from Hacker News

GPT-4o

openai.com

821–830 of 1001 posts

Re: GPT-4o

#821

I found these videos quite hard to watch. There is a level of cringe that I found a bit unpleasant. It’s like some kind of uncanny valley of human interaction that I don’t get on nearly the same level with the text version.

The demo where they take turns singing felt like two nervous slaves trying to please their overlord who kept interrupting them and demanding more harmony.

It is that and that’s okay, they’re algorithms.

Re: GPT-4o

#822
post #24

The most impressive part is that the voice uses the right feelings and tonal language during the presentation. I'm not sure how much of that was that they had tested this over and over, but it is really hard to get that right so if they didn't fake it in some way I'd say that is revolutionary.

That's the part that really struck me. I thought it was particularly impressive with the Sal Khan maths tutor demo and the one with BeMyEyes. The comment at the end about the dog was an interesting ad-lib.

The only slightly annoying thing at the moment is they seem hard to interrupt, which is an important mechanism in conversations. But that seems like a solvable problem. They kind of need to be able to interpret body language a bit to spot when the speaker is about to interrupt.

Re: GPT-4o

#823
Hopefully this will be them turning a new leaf. Making GPT-4 more accessible, cutting API costs, and making a general personal assistant chatbot on iPhone are a lot different than them tracking down and destroying the business of every customer using their API one by one. Let's hope this trend continues.

Re: GPT-4o

#824

Hopefully this will be them turning a new leaf. Making GPT-4 more accessible, cutting API costs, and making a general personal assistant chatbot on iPhone are a lot different than them tracking down and destroying the business of every customer using their API one by one. Let's hope this trend continues.

[flagged]

Re: GPT-4o

#825

This is a very cool demo - if you dig deeper there’s a clip of them having a “blind” AI talk to another AI with live camera input to ask it to explain what it’s seeing. Then they, together, sing a song about what they’re looking at, alternating each line, and rhyming with one another . Given all of the isolated capabilities of AI, this isn’t particularly surprising, but seeing it all work together in real time is pre…

or just... unemployed.

Re: GPT-4o

#826

I worry that this tech will amplify the cultural values we have of "good" and "bad" emotions way more than the default restrictions that social media platforms put on the emoji reactions (e.g., can't be angry on LinkedIn). I worry that the AI will not express anger, not express sadness, not express frustration, not express uncertainty, and many other emotions that the culture of the fine-tuners might believe are "bad…

You did a super job wrapping things up! And I'm not just saying that because I have to!

Re: GPT-4o

#827
post #298

Very interesting and extremely impressive! I tried using the voice chat in their app previously and was disappointed. The big UX problem was that it didn't try to understand when I had finished speaking. English is a second language and I paused a bit too long thinking of a word and it just started responding to my obviously half spoken sentence. Trying again it just became stressful as I had to rush my words out to…

This makes me wonder: if you tell 4o something like "listen to me until I say all done" will it be able to suppress its output until it hears that? I'm guessing not quite possible now, just because I'm guessing patiently waiting is a different band of information that they haven't implemented. But I really don't know.

I'd bet it can't do it now. I'd be curious to hear what it says in response. Partially because it requires time dependent reasoning about being a participant in a conversation.

It shouldn't be too hard to make this work though. If you make the AI start by emitting either a "my turn to talk" or "still listening" token it should be able to listen patiently. If trained correctly.

Re: GPT-4o

#828
post #495

We've had voice input and voice output with computers for a long time, but it's never felt like spoken conversation. At best it's a series of separate voice notes. It feels more like texting than talking. These demos show people talking to artificial intelligence. This is new. Humans are more partial to talking than writing. When people talk to each other (in person or over low-latency audio) there's a rich metadata…

I wonder how it will work in real life and not in a demo…

Besides - not sure if I want this level of immersion/fake when talking to a computer...

"Her" comes to mind pretty quickly…

Re: GPT-4o

#829

Now that I see this, here is my wish (I know there are security privacy concerns but let's pretend there are not there for this wish): An app that runs on my desktop and has access to my screen(s) when I work. At any time I can ask it something about what's on the screen, it can jump in and let me know if it thinks I made a mistake (think pair programming) or a suggestion (drafting a document). It can also quickly ta…

Here you go: UFO - A UI-Focused Agent for Windows OS Interaction

"UFO is a UI-Focused dual-agent framework to fulfill user requests on Windows OS by seamlessly navigating and operating within individual or spanning multiple applications."

https://github.com/microsoft/UFO?tab=readme-ov-file

Re: GPT-4o

#830

For all the hype around this announcement I was expecting more than some demo-level stuff that close to nobody will use in real life. Disappointing.

That's not true, scammers will definitely be using this a lot! Also clueless C-levels who want to nix hundreds of human customer support agents!

You'll get to sit on the phone talking to some convincing robot that won't let you do anything so that the megacorps can save 0.0001 cents! Ain't progress looking so good?

Post reply on HN