Live data from Hacker News

GPT-4.5

openai.com

311–320 of 1001 posts

Re: GPT-4.5

#311
post #167

First impression of GPT-4.5: 1. It is very very slow, for some applications where you want real time interactions is just not viable, the text attached below took 7s to generate with 4o, but 46s with GPT4.5 2. The style it writes is way better: it keeps the tone you ask and makes better improvements on the flow. One of my biggest complaints with 4o is that you want for your content to be more casual and accessible bu…

>1. It is very very slow, ... below took 7s to generate with 4o, but 46s with GPT4.5

This is positively luxurious by o1-pro standards which I'd say average 5 minutes. That said I totally agree even ~45s isn't viable for real-time interactions. I'm sure it'll be optimized.

Of course, my comparing it to the highest-end CoT model in [publicly-known] existence isn't entirely fair since they're sort of apples and oranges.

Re: GPT-4.5

#312

GPT-4.5 Preview scored 45% on aider's polyglot coding benchmark [0]. OpenAI describes it as "good at creative tasks" [1], so perhaps it is not primarily intended for coding. 65% Sonnet 3.7, 32k think tokens (SOTA) 60% Sonnet 3.7, no thinking 48% DeepSeek V3 45% GPT 4.5 Preview [0] https://aider.chat/docs/leaderboards/ [1] https://platform.openai.com/docs/models#gpt-4-5

I was waiting for your comment and wow... that's bad. I guess they are ceding the LLMs for coding market to Anthropic? I remember seeing an industry report somewhere and it claimed software development is the largest user of LLMs, so it seems weird to give up in this area.

I assume they go all in "the new google" direction. Embedded ads coming soon I guess in the free version (chat.com).

Re: GPT-4.5

#313
post #148

Earlier quoted context omitted.

> GPT 4.5 pricing is insane: > I'm still gonna give it a go, though. Seems like the pricing is pretty rational then?

Not if people just try a few prompts then stop using it.

How much of OAI's reported users are doing exactly this?

Re: GPT-4.5

#314
post #303

Earlier quoted context omitted.

Eh, I think o1-pro is by far the most capable model available right now in terms of pure problem solving.

You can try Claude 3.7-Thinking and Grok 3 Think. 10 times cheaper, as good, or very similar to o1-pro.

I haven’t tried Grok yet so can’t speak to that, but I find o1-pro is much stronger than 3.7-thinking for e.g. distributed systems and concurrency problems.

Re: GPT-4.5

#315
post #177

Earlier quoted context omitted.

I’m not an expert or anything, but from my vantage point, each passing release makes Altman’s confidence look more aspirational than visionary, which is a really bad place to be with that kind of money tied up. My financial manager is pretty bullish on tech so I hope he is paying close attention to the way this market space is evolving. He’s good at his job, a nice guy, and surely wears much more expensive underwear…

You're the one buying him the underwear. Don't index funds outperform managed investing? I think especially after accounting for fees, but possibly even after accounting that 50% of money managers are below average.

Not all investing is throwing cash at an index, though. There's other types of investing like direct indexing (to harvest losses), muni bonds, etc.

Paying someone to match your risk profile and financial goals may be worth the fee, which as you pointed out is very measurable. YMMV though.

Re: GPT-4.5

#316

That presentation was super underwhelming. We got to watch them compare… the vibes? … of 4.5 vs o1. No wonder Sam wasn’t part of the presentation.

Sam tweeted "taking care of my kid in the hospital":

https://x.com/sama/status/1895210655944450446

Let's not assume that he's lying. Neither the presentation nor my short usage via the API blew me away, but to really evaluate it, you'd have to use it longer on a daily basis. Maybe that becomes a possiblity with the announced performance optimizations that would lower the price...

Re: GPT-4.5

#317

GPT-2 was laugh out loud funny, rolling on the ground funny. I miss that - newer LLMs seem to have lost their sense of humor. On the other hand GPT-2's funny stories often veered into murdering everyone in the story and committing heinous crimes but that was part of the weird experience.

Totally agree, i think the gargantuan hidden pre prompts, censorship through reinforcement learning and whatever has killed most creativity.

The newer models are incredible, but the tone is just soul sucking even when it tries to be "looser" in the later iterations.

Re: GPT-4.5

#319
post #167

First impression of GPT-4.5: 1. It is very very slow, for some applications where you want real time interactions is just not viable, the text attached below took 7s to generate with 4o, but 46s with GPT4.5 2. The style it writes is way better: it keeps the tone you ask and makes better improvements on the flow. One of my biggest complaints with 4o is that you want for your content to be more casual and accessible bu…

I opened your link in a new tab and looked at it a couple minutes later. By then I forgot which was o and which was .5 I honestly couldn't decide which I prefer

I definitely prefer the 4.5, but that might just be because it sounds 'less like ChatGPT', ironically.

Re: GPT-4.5

#320
post #9

One comparison I found interesting... I think GPT-4o has a more balanced answer! > What are your thoughts on space exploration? GPT-4.5: Space exploration isn't just valuable—it's essential. People often frame it as a luxury we pursue after solving Earth-bound problems. But space exploration actually helps us address those very challenges: climate change (via satellite monitoring), resource scarcity (through asteroid…

Indeed, and the difference could in essence be achieved yourself with a different system prompt on 4o. What exactly is 4.5 contributing here in terms of a more nuanced intelligence?

The new RLHF direction (heavily amplified through scaling synthetic training tokens) seems to clobber any minor gains the improved base internet prediction gains might've added.

Post reply on HN