Live data from Hacker News

GPT-4o

openai.com

671–680 of 1001 posts

Re: GPT-4o

#671
This is the first demo where you can really sense that beating LLM benchmarks should not be the target. Just remember the time when the iPhone has meager specs but ultimately delivered a better phone experience than the competition.

This is the power of the model where you can own the whole stack and build a product. Open Source will focus on LLM benchmarks since that is the only way foundational models can differentiate themselves, but it does not mean it is a path to a great user experience.

So Open Source models like Llama will be here to stay, but it feels more like if you want to build a compelling product, you have to own and control your own model.

Re: GPT-4o

#672
post #287
post #24

The most impressive part is that the voice uses the right feelings and tonal language during the presentation. I'm not sure how much of that was that they had tested this over and over, but it is really hard to get that right so if they didn't fake it in some way I'd say that is revolutionary.

Crazy that interruption also seems to work pretty smoothly

Really? I think interruption and timing in general still seems like a problem that has yet to be solved. It was the most janky aspect of the demos imo.

Re: GPT-4o

#673
post #598
post #495

We've had voice input and voice output with computers for a long time, but it's never felt like spoken conversation. At best it's a series of separate voice notes. It feels more like texting than talking. These demos show people talking to artificial intelligence. This is new. Humans are more partial to talking than writing. When people talk to each other (in person or over low-latency audio) there's a rich metadata…

im human and much much more partial to typing than talking. talking is a lot of work for me and i can't process my thinking well at all without writing.

The good news is the interface will be multi modal. Talk, type, and I guess someday just think.

Re: GPT-4o

#674

I found these videos quite hard to watch. There is a level of cringe that I found a bit unpleasant. It’s like some kind of uncanny valley of human interaction that I don’t get on nearly the same level with the text version.

All these demo style ads/videos are super jarring and uncanny valley-esque to watch as an Australian. The US corporate cultural norms are super bizarre to the rest of the world, and the California based holy omega of tech companies really takes this to the extreme. The application might work well if you interact with it like you are a normal human being - but I can't tell because this presentation is corporate robots…

[flagged]

Re: GPT-4o

#675

I’m really not impressed. My academic background is in a field where there are lots of public misconceptions. It does an absolutely terrible job. Even basic textbook things where there isn’t much public misunderstanding are “rounded” to what sounds smart.

What field? Curious to see it myself

Re: GPT-4o

#676

I've worked quite a bit with STT and TTS over the past ~7 years, and this is the most impressive and even startling demo I've seen. But I would like to see how this is integrated into applications by third party developers where the AI is doing a specific job. Is it still as impressive? The biggest challenge I've had with building any autonomous "agents" with generic LLM's is they are overly gullible and accommodatin…

>Also STT is rife with speaker interjections, leading to significant user frustrations and they just want to talk to a person. Hard to see if this is really solved yet.

This is not using TTS or STT. Audio and Image data can be tokenized as readily as text. This is simply a LLM that happens to have been trained to receive and spit out audio and image tokens as well as text tokens. Interjections are a lot more palatable in this paradigm as most of the demos show.

Re: GPT-4o

#677
post #495

We've had voice input and voice output with computers for a long time, but it's never felt like spoken conversation. At best it's a series of separate voice notes. It feels more like texting than talking. These demos show people talking to artificial intelligence. This is new. Humans are more partial to talking than writing. When people talk to each other (in person or over low-latency audio) there's a rich metadata…

> Humans are more partial to talking than writing. Amazon, Google, and Apple have sunk literally billions of dollars into this idea only to find out that, no, we aren't. We are with other humans, yes. When socialization is part of the conversation. When I'm talking to my local barista I'm not just ordering a coffee, I'm also maintaining a relationship with someone in my community. But when it comes to work, writing >…

Writing is only superior to conversation when weighed against discussions with more than 3 people. A quick call with one or two other people always results in more progress being made as long as everyone involved wants to get it done. Messaging back and forth takes much more time and often leads to misunderstandings.

Re: GPT-4o

#678

Won't this make pretty much all of the work to make a website accessible go away, as it becomes cheap enough? Why struggle to build alt content for the impaired when it can be generated just in time as needed? And much the same for internationalization.

Because accessibility is more than checking a box. Got a photo you took on your website? Alternative text that you wrote capturing what you see in the photo you took is accessibility done right. Alt text generated by bots is not accessibility done right unless that bot knows what you see in the photo you took and that's not likely to happen.

Re: GPT-4o

#679
post #24

The most impressive part is that the voice uses the right feelings and tonal language during the presentation. I'm not sure how much of that was that they had tested this over and over, but it is really hard to get that right so if they didn't fake it in some way I'd say that is revolutionary.

I mention this down thread, but a symptom of a tech product of sufficient advancement is the nature of its introduction matters less and less. Based on the casual production of these videos, the product must be this good. https://news.ycombinator.com/item?id=40346002

I noticed this as well. They gave zero fs about the fit and finish of these videos because they know this is magic in a bottle.

Re: GPT-4o

#680
post #569
post #495

We've had voice input and voice output with computers for a long time, but it's never felt like spoken conversation. At best it's a series of separate voice notes. It feels more like texting than talking. These demos show people talking to artificial intelligence. This is new. Humans are more partial to talking than writing. When people talk to each other (in person or over low-latency audio) there's a rich metadata…

But in this case you're not talking with a real person. Instinctively, I dislike a robot that pretends to be a real human being.

[deleted]
Post reply on HN