Live data from Hacker News

GPT-4o

openai.com

191–200 of 1001 posts

Re: GPT-4o

#191

This thing continues to stress my skepticism for AI scaling laws and the broad AI semiconductor capex spending. 1- OpenAI is still working in GPT-4-level models. More than 14 months after the launch of GPT-4 and after more than $10B in capital raised. 2- The rhythm that token prices are collapsing is bizarre. Now a (bit) better model for 50% of the price. How people seriously expect these foundational model companies…

Tbf gpt4 level seems useful and better than almost everything else (or close if not). The more important barriers for use in applications have been cost, throughout and latency. Oh and modalities, which have expanded hugely.

Re: GPT-4o

#192
It's pretty impressive, although I don't like the voice / tone, I prefer something more neutral.

Re: GPT-4o

#193
post #43
post #24

The most impressive part is that the voice uses the right feelings and tonal language during the presentation. I'm not sure how much of that was that they had tested this over and over, but it is really hard to get that right so if they didn't fake it in some way I'd say that is revolutionary.

(I work at OpenAI.) It's really how it works.

How far are we away from something like a helmet with chat GPT and a video camera installed, I imagine this will be awesome for low vision people. Imagine having a guide tell you how to walk to the grocery store, and help you grocery shop without an assistant. Of course you have tons of liability issues here, but this is very impressive

Re: GPT-4o

#194
GPT-4o being a truly multimodal model is exciting, does open the door to more interesting products. I was curious about the new tokenizer which uses much fewer tokens for non-English, but also 1.1x fewer tokens for English, so I'm wondering if this means each token now can be more possible values than before? Might make sense provided that they now also have audio and image output tokens? https://openai.com/index/hello-gpt-4o/

I wonder what "fewer tokens" really means then, without context on raising the size of each token? It's a bit like saying my JPEG image is now using 2x fewer words after I switched from a 32-bit to a 64-bit architecture no?

Re: GPT-4o

#195

This thing continues to stress my skepticism for AI scaling laws and the broad AI semiconductor capex spending. 1- OpenAI is still working in GPT-4-level models. More than 14 months after the launch of GPT-4 and after more than $10B in capital raised. 2- The rhythm that token prices are collapsing is bizarre. Now a (bit) better model for 50% of the price. How people seriously expect these foundational model companies…

>Token volume needs to double just for revenue to stand still

Profits are the real metric. Token volume doesn't need to double for profits to stand still if operational costs go down.

Re: GPT-4o

#196
post #24

The most impressive part is that the voice uses the right feelings and tonal language during the presentation. I'm not sure how much of that was that they had tested this over and over, but it is really hard to get that right so if they didn't fake it in some way I'd say that is revolutionary.

Right to who? To me, the voice sounds like an over enthusiastic podcast interviewer. Whats wrong with wanting computers to sound like what people think computers should sound like?

It understands tonal language, you can tell it how you want it to talk, I have never seen a model like that before. If you want it to talk like a computer you can tell it to, they did it during the presentation, that is so much better than the old attempts at solving this.

Re: GPT-4o

#197
post #43
post #24

The most impressive part is that the voice uses the right feelings and tonal language during the presentation. I'm not sure how much of that was that they had tested this over and over, but it is really hard to get that right so if they didn't fake it in some way I'd say that is revolutionary.

(I work at OpenAI.) It's really how it works.

Is it possible to use this as a TTS model? I noticed on the announcement post that this is a single model as opposed to a text model being piped to a separate TTS model.

Re: GPT-4o

#200

This is really impressive engineering. I thought real time agents would completely change the way we're going to interact with large models but it would take 1~2 more years. I wonder what kind of new techs are developed to enable this, but OpenAI is fairly secretive so we won't be able to know their sauce. On the other hand, this also feels like a signal that reasoning capability has probably already been plateaued a…

Why would reasoning have plateau’d?

I think reasoning ability is not the largest bottleneck for improvement in usefulness right now. Cost is a bigger one IMO.

Running these models as agents is hella expensive, and agents or agent-like recurrent reasoning (like humans do) is the key to improved performance if you look at any type of human intelligence.

Single-shot performance only gets you so far.

For example- If it can write code 90% of the way, and then debug in a loop, it’d be much more performant than any single shot algorithm.

And OpenAI has these huge models in their basement probably. But they might not be much more useful than GPT-4 when used as single-shot. I mean, what could it do what we can’t do today with gpt-4?

It’s agents and recurrent reasoning we need for more usefulness.

At least- That’s my humble opinion as an amateur neuroscientist that plays around with these models.

Post reply on HN