Live data from Hacker News

GPT-4o

openai.com

321–330 of 1001 posts

Re: GPT-4o

#321
post #115

The usual critics will quickly point out that LLMs like GPT-4o still have a lot of failure modes and suffer from issues that remain unresolved. They will point out that we're reaping diminishing returns from Transformers. They will question the absence of a "GPT-5" model. And so on -- blah, blah, blah, stochastic parrots, blah, blah, blah. Ignore the critics. Watch the demos. Play with it. This stuff feels magical .…

Magical?

the interruptiopn part is just flow control at the edge. control-s, control-c stuff, right? not AI?

The sound of a female voice to an audience 85% composed of males between the ages of 14 and 55 is "magical", not this thing that recreates it.

so yeah, its flow control and compression of highly curated, subtle soft porn. Subtle, hyper targeted, subconscious porn honed by the most colossal digitally mediated focus group ever constructed to manipulate our (straight male) emotions.

why isn't the voice actually the voice of the pissed off high school janitor telling you to man-up and stop hyperventilating? instead its a woman stroking your ego and telling you to relax and take deep breaths. what dataset did they train that voice on anyway?

Re: GPT-4o

#322

This is really impressive engineering. I thought real time agents would completely change the way we're going to interact with large models but it would take 1~2 more years. I wonder what kind of new techs are developed to enable this, but OpenAI is fairly secretive so we won't be able to know their sauce. On the other hand, this also feels like a signal that reasoning capability has probably already been plateaued a…

Ya so sad that OpenAI isn't more Open imagine if OpenAI was still sharing their thought processes and papers with the overall commity, really wish we saw collaborations between OpenAI and Meta for instance to really have helped push the open source arena further ahead, i love that their latest models are so great but the fact they aren't helping the Open source arena to progress is sad. Imagine how far we'd be if OpenAI was still as open as they once were and we saw collaborations betweeen Meta, OpenAI and Anthropic all working and sharing growth and tech to reduce double work and help each other not go down failed paths.

Re: GPT-4o

#323
post #194

GPT-4o being a truly multimodal model is exciting, does open the door to more interesting products. I was curious about the new tokenizer which uses much fewer tokens for non-English, but also 1.1x fewer tokens for English, so I'm wondering if this means each token now can be more possible values than before? Might make sense provided that they now also have audio and image output tokens? https://openai.com/index/hel…

The size can stay the same. Tokens get converted into state which is a vector of 4000+ dimensions. So you could have millions of tokens even and still encode them into the same state size.

Re: GPT-4o

#324

This thing continues to stress my skepticism for AI scaling laws and the broad AI semiconductor capex spending. 1- OpenAI is still working in GPT-4-level models. More than 14 months after the launch of GPT-4 and after more than $10B in capital raised. 2- The rhythm that token prices are collapsing is bizarre. Now a (bit) better model for 50% of the price. How people seriously expect these foundational model companies…

> where are my agents?

https://github.com/OpenAdaptAI/OpenAdapt/

Re: GPT-4o

#325

Big questions are (1) when is this going to be rolled out to paid users? (2) what is the remaining benefit of being a paid user if this is rolled out to free users? (3) Biggest concern is will this degrade the paid experience since GPT-4 interactions are already rate limited. Does OpenAI have the hardware to handle this? Edit: according to @gdb this is coming in "weeks" https://twitter.com/gdb/status/1790074041614717…

This might mean GPT-5 is coming soon and it will only be available to paid users.

Or they just made a bunch of money on their licensing deal with Apple. So they don't need to charge for ChatGPT anymore.

Re: GPT-4o

#326
post #312

I cannot believe that that overly excited giggle tone of voice you see in the demo videos made it through quality control?! I've only watched two videos so far and it's already annoying me to the point that I couldn't imagine using it regularly.

[deleted]

Re: GPT-4o

#327
I’m a huge user of GPT4 and Opus in my work but I’m a huge user of GPT4-Turbo voice in my personal life. I use it on my commutes to learn all sorts of stuff. I’ve never understood the details of cameras and the relationship between shutter speed and aperture and iso in a modern dslr which given the aurora was important. We talked through and I got to an understanding in a way having read manuals and textbooks didn’t really help before. I’m a much better learner by being able to talk and hear and ask questions and get responses.

Extend this to quantum foam, to ergodic processes, to entropic force, to Darius and Xerces, to poets of the 19th century - it’s changed my life. Really glad to see an investment in stream lining this flow.

Re: GPT-4o

#328
Just something I noticed in the Language tokenization section

When referring to itself, it uses the female word in Marathi नमस्कार, माझे नाव जीपीटी-4o आहे| मी एक नवीन प्रकारची भाषा मॉडेल आहे| तुम्हाला भेटून आनंद झाला!

and Male word in Hindi नमस्ते, मेरा नाम जीपीटी-4o है। मैं एक नए प्रकार का भाषा मॉडल हूँ। आपसे मिलकर अच्छा लगा!

Re: GPT-4o

#329

This is really impressive engineering. I thought real time agents would completely change the way we're going to interact with large models but it would take 1~2 more years. I wonder what kind of new techs are developed to enable this, but OpenAI is fairly secretive so we won't be able to know their sauce. On the other hand, this also feels like a signal that reasoning capability has probably already been plateaued a…

This isn't really new tech, it's just an async agent in front of a multimodal model. It seems from the demo that the improvements have been in response latency and audio generation. Still, it looks like they're building a solid product, which has been their big issue so far.

Its 200-300ms for a multimodal response, thats REALLY a big step forward, especially given it's doing it with full voice response, not just text.
Post reply on HN