Live data from Hacker News

GPT-4.5

openai.com

11–20 of 1001 posts

Re: GPT-4.5

#11
I am beginning to think these human eval tests are a waste of time at best, and negative value at worst. Maybe I am being snobby, but I don't think the average human is able to properly evaluate usefulness, truthfulness, or other metrics that I actually care about. I am sure this is good for openAI since if more people like what the hear, they are more likely come back.

I don't want my AI more obsequious, I want it more correct and capable.

My only use case is coding though, so maybe I am not representative of their usual customers?

Re: GPT-4.5

#12
> Because of this, we’re evaluating whether to continue serving it in the API long-term as we balance supporting current capabilities with building future models.

Seems like it's not going to be deployed for long.

$75.00 / 1M tokens for input

$150.00 / 1M tokens for output

That's crazy prices.

Re: GPT-4.5

#14
Oh this makes sense. chatGPT results have taken a nose dive in quality lately.

It couldn't write a simple rename function for me yesterday, still buggy after seven attempts.

I'm more and more convinced that they dumb down the core product when they plan to release a new version to make the difference seem bigger.

Re: GPT-4.5

#16
Sounds like it's a distill of O1? After R1, I don't care that much about non-reasoning models anymore. They don't even seem excited about it on the livestream.

I want tiny, fast and cheap non-reasoning models I can use in APIs and I want ultra smart reasoning models that I can query a few times a day as an end user (I don't mind if it takes a few minutes while I refill a coffee).

Oh, and I want that advanced voice mode that's good enough at transcription to serve as a babelfish!

After that, I guess it's pretty much all solved until the robots start appearing in public.

Re: GPT-4.5

#17

Oh this makes sense. chatGPT results have taken a nose dive in quality lately. It couldn't write a simple rename function for me yesterday, still buggy after seven attempts. I'm more and more convinced that they dumb down the core product when they plan to release a new version to make the difference seem bigger.

99% chance that's confirmation bias

Re: GPT-4.5

#19

I am beginning to think these human eval tests are a waste of time at best, and negative value at worst. Maybe I am being snobby, but I don't think the average human is able to properly evaluate usefulness, truthfulness, or other metrics that I actually care about. I am sure this is good for openAI since if more people like what the hear, they are more likely come back. I don't want my AI more obsequious, I want it m…

The SuperTuring era.
Post reply on HN