Live data from Hacker News

GPT-4o

openai.com

291–300 of 1001 posts

Re: GPT-4o

#291
after watching the OpenAI videos I'm looking at my sad Google Assistant speaker in the corner.

Come on Google... you can update it.

Re: GPT-4o

#292

I admit I drink the koolaid and love LLMs and their applications. But damn, the way it’s responds in the demo gave me goosebumps in a bad way. Like an uncanny valley instincts kicks in.

The chuckling made me uneasy for some reason lol. Calm down, you're not like us. Don't pretend!

Can't wait for Meta's version 2 years down the line that someone will eventually fine tune to Agent Smith's personality and voice.

"Evolution, human. Evolution. Like the dinosaur. Look out that window. You've had your time. The future is our world. The future is our time."

Re: GPT-4o

#293

This thing continues to stress my skepticism for AI scaling laws and the broad AI semiconductor capex spending. 1- OpenAI is still working in GPT-4-level models. More than 14 months after the launch of GPT-4 and after more than $10B in capital raised. 2- The rhythm that token prices are collapsing is bizarre. Now a (bit) better model for 50% of the price. How people seriously expect these foundational model companies…

Where I work in the hoary fringes of high end tech we can’t secure enough token processing for our use cases. Token price decreases means opening of capacity but we immediately hit the boundaries of what we can acquire. We can’t keep up with the use cases - but more than that we can’t develop tooling to harness things fast enough and the tooling we are creating is a quick hack. I don’t fear for the revenue of base model providers. But I think in the end the person selling the tools makes the most and in this case I think it continue to be cloud providers. I think in a very real way OpenAI and Anthropic are commercialized charities driving change and commoditizing rapidly their own products and it’ll be infrastructure providers who win the high end model game. I don’t think this is a problem I think this is in fact inline with their original charters but a different path than most people view nonprofit work. A much more capitalist and accelerated take.

Where they might make future businesses is in the tooling. My understanding from friends within these companies is their tooling is remarkably advanced vs generally available tech. But base models aren’t the future of revenues (to be clear tho they make considerable revenue today but at some point their efficiency will cannibalize demand and the residual business will be tools)

Re: GPT-4o

#294
I was about to say how this thing is lame because it sounds so forced and robotic and fake, and even though the intonations do make it sound more human-like, it's very clear that they made a big effort to make it sound like natural speech, but failed.

...but then I realized that's basically the kind of thing Data from Star Trek struggles with as part of his character. We're almost in that future, and I'm already falling into the role of the ignorant human that doesn't respect androids.

Re: GPT-4o

#295
I think people excited should look at the empty half of the glass here, this is pretty much an admitance that they are struggling to go past gpt 4 on a significant scale.

Not like they have to be scared yet, I mean Google has yet to release their vaporware Ultra model that is supposedly like 1% better than GPT 4 in some metrics...

I smell an AI crash coming in a few years if they can't actually get this stuff usable for day to day life.

Re: GPT-4o

#296
post #27

Parts of the demo were quite choppy (latency?) so this definitely feels rushed in response to Google I/O. Other than that, looks good. Desktop app is great, but I didn’t see no mention of being able to use your own API key so OS projects might still be needed. The biggest thing is bringing GPT-4 to free users, that is an interesting move. Depending on what the limits are, I might cancel my subscription.

They need to fade the audio or add some vocal queue when it's being interrupted. It makes it sound like it's losing connection. What'll be really impressive is when it intentionally starts interrupting you.

Agree. While watching the demo video, I thought I was the one having connectivity issues.

Re: GPT-4o

#297
So far, I'm impressed. It seems to be significantly better than GPT-4 at accessing current online documentation and forming answers that use it effectively. I've been asking it to do so, and it has.

Re: GPT-4o

#298
Very interesting and extremely impressive!

I tried using the voice chat in their app previously and was disappointed. The big UX problem was that it didn't try to understand when I had finished speaking. English is a second language and I paused a bit too long thinking of a word and it just started responding to my obviously half spoken sentence. Trying again it just became stressful as I had to rush my words out to avoid an annoying response to an unfinished thought.

I didn't try interrupting it but judging by the comments here it was not possible.

It was very surprising to me to be so overtly exposed to the nuances of real conversation. Just this one thing of not understanding when it's your turn to talk made the interaction very unpleasant, more than I would have expected.

On that note, I noticed that the AI in the demo seems to be very rambly. It almost always just kept talking and many statements were reiterations of previous ones. It reminded me of a type of youtuber that uses a lot of filler phrases like "let's go ahead and ...", just to be more verbose and lessen silences.

Most of the statements by the guy doing the demo were interrupting the AI.

It's still extremely impressive but I found this interesting enough to share. It will be exciting to see how hard it is to reproduce these abilities in the open, and to solve this issue.

Re: GPT-4o

#299

OAI just made an embarrassment of Google's fake demo earlier this year. Given how this was recorded, I am pretty certain it's authentic.

This feature has been in iOS for a while now, just really slow and without some of the new vision aspects. This seems like a version 2 for me.

That old feature uses Whisper to transcribe your voice to text, and then feeds the text into the GPT which generates a text response, and then some other model synthesizes audio from that text.

This new feature feeds your voice directly into the GPT and audio out of it. It’s amazing because now ChatGPT can truly communicate with you via audio instead of talking through transcripts.

New models should be able to understand and use tone, volume, and subtle cues when communicating.

I suppose to an end user it is just “version 2” but progress will become more apparent as the natural conversation abilities evolve.

Re: GPT-4o

#300
it seems like the ability to interrupt is more like the interrupt in the computer sense ... A control-c (or control-s tty flow control for you old timers), not a cognitive evaluation followed by the "reasoned" decision to pause voice output. not that it matters i guess, its just not general intelligence. its just flow control.

but also, thats why it fails a real turing test. a real person would be irritated as fuck by the interruptions

Post reply on HN