Live data from Hacker News

GPT-4.5

openai.com

351–360 of 1001 posts

Re: GPT-4.5

#351
This looks like a first generation model to bootstrap future models from, not a competitive product at all. The knowledge cutoff is pretty old as well. (2023, seriously?)

If they wanted to train it to have some character like Anthropic did with Claude 3... honestly I'm not seeing it, at least not in this iteration. Claude 3 was/is much much more engaging.

Re: GPT-4.5

#352
post #309
post #280

Earlier quoted context omitted.

It hallucinates at 37% on SimpleQA yeah, which is a set of very difficult questions inviting hallucinations. Claude 3.5 Sonnet (the June 2024 editiom, before October update and before 3.7) hallucinated at 35%. I think this is more of an indication of how behind OpenAI has been in this area.

Are the benchmarks known ahead of time? Could the answer to the benchmarks be in the training data?

In general yes, bench mark pollution is a big problem and why only dynamic benchmarks matter.

Re: GPT-4.5

#353

Earlier quoted context omitted.

> "Early testing shows that interacting with GPT‑4.5 feels more natural. Its broader knowledge base, improved ability to follow user intent, and greater “EQ” make it useful for tasks like improving writing, programming, and solving practical problems. We also expect it to hallucinate less." "Early testing doesn't show that it hallucinates less, but we expect that putting that sentence nearby will lead you to draw a c…

The usage of "greater" is also interesting. It's like they are trying to say better, but greater is a geographic term and doesn't mean "better" instead it's closer to "wider" or "covers more area."

[deleted]

Re: GPT-4.5

#354

It is interesting that they are focusing a large part of this release on the model having a higher "EQ" (Emotional Quotient). We're far from the days of "this is not a person, we do not want to make it addictive" and getting a firm foot on the territory of "here's your new AI friend". This is very visible in the example comparing 4o with 4.5 when the user is complaining about failing a test, where 4o's response is wh…

> We're far from the days of "this is not a person, we do not want to make it addictive" and getting a firm foot on the territory of "here's your new AI friend".

That’s a hard nope from me, when companies pull that move. I’ll stick to my flesh and blood humans who still hallucinate but only rarely.

Re: GPT-4.5

#355
post #26

GPT 4.5 pricing is insane: Price Input: $75.00 / 1M tokens Cached input: $37.50 / 1M tokens Output: $150.00 / 1M tokens GPT 4o pricing for comparison: Price Input: $2.50 / 1M tokens Cached input: $1.25 / 1M tokens Output: $10.00 / 1M tokens It sounds like it's so expensive and the difference in usefulness is so lacking(?) they're not even gonna keep serving it in the API for long: > GPT‑4.5 is a very large and comput…

Now the real question about AI automation starts. Is it cheaper to pay a human to do the task or a AI company?

Humans have all sorts of issues you have to deal with. Being hungover, not sleeping well, having a personality, being late to work, not being able to work 24/7, very limited ability to copy them. If there's a soulless generic office-droidGPT that companies could hire that would never talk back and would do all sorts of menial work without needing breaks or to use the bathroom, I don't know that we humans stand a chance!

I have a bunch of work that needs doing. I can do it myself, or I can hire one person to do it. I gotta train them and manage them and even after I train them theres still only going to be one of them, and it's subject to their availability. On the other hand, if I need to train an AI to do it, but I can copy that AI, and then spin them up/down like on demand computer in the cloud, and not feel remotely bad about spinning them down?

It's definitely not there yet, but it's not hard to see the business case for it.

Re: GPT-4.5

#356
post #218

Earlier quoted context omitted.

"I knew the dame was trouble the moment she walked into my office." "Uh... excuse me, Detective Nick Danger? I'd like to retain your services." "I waited for her to get the the point." "Detective, who are you talking to?" "I didn't want to deal with a client that was hearing voices, but money was tight and the rent was due. I pondered my next move." "Mr. Danger, are you... narrating out loud?" "Damn! My internal chai…

I don't know if it was you or someone else who made pretty much the same point a few days ago. But I still like it. It makes the whole thing a lot more fun.

https://news.ycombinator.com/context?id=43118925

I've been banging that particular drum for a while on HN, and the mental-model still feels so intuitively strong to me that I'm starting to have doubts: "It feels too right, I must be wrong in some subtle yet devastating way."

Re: GPT-4.5

#357

It is interesting that they are focusing a large part of this release on the model having a higher "EQ" (Emotional Quotient). We're far from the days of "this is not a person, we do not want to make it addictive" and getting a firm foot on the territory of "here's your new AI friend". This is very visible in the example comparing 4o with 4.5 when the user is complaining about failing a test, where 4o's response is wh…

Anthropic pretty much abandoned this direction after Claude 3, and said it wasn't what they wanted [1]. Claude 3.5+ is extremely dry and neutral, it doesn't seem to have the same training.

>Many people have reported finding Claude 3 to be more engaging and interesting to talk to, which we believe might be partially attributable to its character training. This wasn’t the core goal of character training, however. Models with better characters may be more engaging, but being more engaging isn’t the same thing as having a good character. In fact, an excessive desire to be engaging seems like an undesirable character trait for a model to have.

[1] https://www.anthropic.com/research/claude-character

Re: GPT-4.5

#358
post #270
post #167

First impression of GPT-4.5: 1. It is very very slow, for some applications where you want real time interactions is just not viable, the text attached below took 7s to generate with 4o, but 46s with GPT4.5 2. The style it writes is way better: it keeps the tone you ask and makes better improvements on the flow. One of my biggest complaints with 4o is that you want for your content to be more casual and accessible bu…

> It is very very slow Could that be partially due to a big spike in demand at launch?

[deleted]

Re: GPT-4.5

#359

Earlier quoted context omitted.

I suppose this was their final hurrah after two failed attempts at training GPT-5 with the traditional pre-training paradigm. Just confirms reasoning models are the only way forward.

What it confirms, I think, is, that we are going to need a lot more chips.

Further confirmation, IMO, that the idea that any of this leads to anything close to AGI is people getting high on their own supply (in some cases literally).

LLMs are a great tool for what is effectively collected knowledge search and summary (so long as you are willing to accept that you have to verify all of the 'knowledge' they spit back because they always have the ability to go off the rails) but they have been hitting the limits on how much better that can get without somehow introducing more real knowledge for close to 2 years now and everything since then is super incremental and IME mostly just benchmark gains and hype as opposed to actually being purely better.

I personally don't believe that more GPUs solves this, like, at all. But its great for Nvidia's stock price.

Re: GPT-4.5

#360
post #167

First impression of GPT-4.5: 1. It is very very slow, for some applications where you want real time interactions is just not viable, the text attached below took 7s to generate with 4o, but 46s with GPT4.5 2. The style it writes is way better: it keeps the tone you ask and makes better improvements on the flow. One of my biggest complaints with 4o is that you want for your content to be more casual and accessible bu…

How does it compare with o1 and o3 preview?

o3 is okay for text checking but has issues following the prompt correctly, same as o1 and DeepSeek R1, I feel that I need to prompt smaller snippets with them.

Here is the o3 vs a new run of the same text in GPT 4.5

https://www.diffchecker.com/ZEUQ92u7/

Post reply on HN