Live data from Hacker News

GPT-4o

openai.com

241–250 of 1001 posts

Re: GPT-4o

#242
post #58

They are admitting[1] that the new model is the gpt2-chatbot that we have seen before[2]. As many highlighted there, the model is not an improvement like GPT3->GPT4. I tested a bunch of programming stuff and it was not that much better. It's interesting that OpenAI is highlighting the Elo score instead of showing results for many many benchmarks that all models are stuck at 50-70% success. [1] https://twitter.com/Lia…

I think the live demo that happened on the livestream is best to get a feel for this model[0]. I don't really care whether it's stronger than gpt-4-turbo or not. The direct real-time video and audio capabilities are absolutely magical and stunning . The responses in voice mode are now instantaneous, you can interrupt the model, you can talk to it while showing it a video, and it understands (and uses) intonation and…

Hectic!

Thanks for this.

Re: GPT-4o

#243

This thing continues to stress my skepticism for AI scaling laws and the broad AI semiconductor capex spending. 1- OpenAI is still working in GPT-4-level models. More than 14 months after the launch of GPT-4 and after more than $10B in capital raised. 2- The rhythm that token prices are collapsing is bizarre. Now a (bit) better model for 50% of the price. How people seriously expect these foundational model companies…

what do you actually expect from an "agent"?

Re: GPT-4o

#244
post #72

In the first video the AI seems excessively chatty.

Yes, it sounds like an awkwardly perky and over-chatty telemarketer that really wants to be your friend. I find the tone maximally annoying and think most users will find it both stupid and creepy. Based on user preferences, I expect future interactive chat AIs will default to an engagement mode that's optimized for accuracy and is both time-efficient and cognitively efficient for the user.

I suspect this AI Human engagement style will evolve over time to become quite unlike human to human engagement, probably mixing speech with short tones for standard responses like "understood", "will do", "standing by" or "need more input". In the future these old-time demo videos where an AI is forced to do a creepy caricature of an awkward, inauthentic human will be embarrassingly retro-cringe. "Okay, let's do it!"

Re: GPT-4o

#245
post #58

They are admitting[1] that the new model is the gpt2-chatbot that we have seen before[2]. As many highlighted there, the model is not an improvement like GPT3->GPT4. I tested a bunch of programming stuff and it was not that much better. It's interesting that OpenAI is highlighting the Elo score instead of showing results for many many benchmarks that all models are stuck at 50-70% success. [1] https://twitter.com/Lia…

> As many highlighted there, the model is not an improvement like GPT3->GPT4. The improvements they seem to be hyping are in multimodality and speed (also price – half that of GPT-4 Turbo – though that’s their choice and could be promotional, but I expect it’s at least in part, like speed, a consequence of greater efficiency), not so much producing better output for the same pure-text inputs.

the model scores 60 points higher in lmsys than the best gpt 4 turbo model from april, that's still a pretty significant jump in text capability

Re: GPT-4o

#246
post #24

The most impressive part is that the voice uses the right feelings and tonal language during the presentation. I'm not sure how much of that was that they had tested this over and over, but it is really hard to get that right so if they didn't fake it in some way I'd say that is revolutionary.

Right to who? To me, the voice sounds like an over enthusiastic podcast interviewer. Whats wrong with wanting computers to sound like what people think computers should sound like?

Right... enthusiastic and generally confused. It's uncanny valley level expressions. Still better than drab, monotonous speech though.

Re: GPT-4o

#247

This is really impressive engineering. I thought real time agents would completely change the way we're going to interact with large models but it would take 1~2 more years. I wonder what kind of new techs are developed to enable this, but OpenAI is fairly secretive so we won't be able to know their sauce. On the other hand, this also feels like a signal that reasoning capability has probably already been plateaued a…

This isn't really new tech, it's just an async agent in front of a multimodal model. It seems from the demo that the improvements have been in response latency and audio generation. Still, it looks like they're building a solid product, which has been their big issue so far.

Re: GPT-4o

#248
post #194

GPT-4o being a truly multimodal model is exciting, does open the door to more interesting products. I was curious about the new tokenizer which uses much fewer tokens for non-English, but also 1.1x fewer tokens for English, so I'm wondering if this means each token now can be more possible values than before? Might make sense provided that they now also have audio and image output tokens? https://openai.com/index/hel…

New tokenizer has a much larger vocabulary (200k)[0].

[0] https://github.com/openai/tiktoken/commit/9d01e5670ff50eb74c...

Re: GPT-4o

#249
In my experience so far, GPT-4o seems to sit somewhere between the capability of GPT-3.5 and GPT-4.

I'm working on an app that relies more on GPT-4's reasoning abilities than inference speed. For my use case, GPT-4o seems to do worse than GPT-4 Turbo on reasoning tasks. For me this seems like a step-up from GPT-3.5 but not from GPT-4 Turbo.

At half the cost and significantly faster inference speed, I'm sure this is a good tradeoff for other use cases though.

Re: GPT-4o

#250

I found these videos quite hard to watch. There is a level of cringe that I found a bit unpleasant. It’s like some kind of uncanny valley of human interaction that I don’t get on nearly the same level with the text version.

While it is probably pretty normal for California, the insincere flattery and patronizing eagerness are definitely grating But then you have to stack that up against the fact that we are examining a technology and nitpicking over its tone of voice.
Post reply on HN