Live data from Hacker News

GPT-4.5

openai.com

811–820 of 1001 posts

Re: GPT-4.5

#811
post #114

Who wants a model that is not reasoning? The older models are just fine.

They said this is their last non reasoning model so I'm assuming there is a sunk cost aspect to it.

Re: GPT-4.5

#812

Is it official then? Most of us have been waiting for this moment for a while. The transformer architecture as it is currently understood can't be milked any further. Many of us knew this since last year. GPT-5 delays eventually led to non-tech voices to suggest likewise. But we all held our final decision until the next big release from OpenAI as Sam Altman has been making claims about AGI entering the workforce thi…

> will nonetheless make people's lives better While I mostly agree with your assessment, I am still not convinced of this part. Right now, it may be making our lives marginally better. But once the enshittification starts to set in, I think it has the potential to make things a lot worse. E.g. I think the advertisement industry will just love the idea of product placements and whatnots into the AI assistant conversat…

*good*. the answer to this is legislation —- legally, stop allowing shitty ads everywhere all the time. I hope these problems we already have are exacerbated by the ease of generating content with LLMs and people actually have to think for themselves again

Re: GPT-4.5

#813
Funny times. Sonnet 3.7 launches and there is big hype... but complaints start to surface on r/cursor that it is doing too much, is too confident, has no personality. I wonder if 4.5 will be the reverse, an under-hyped launch, but a dawning realisation that it is incredibly useful. Time will tell!

Re: GPT-4.5

#814

Earlier quoted context omitted.

I'd put myself on the pessimistic side of all the hype, but I still acknowledge that where we are now is a pretty staggering leap from two years ago. Coding in particular has gone from hints and fragments to full scripts that you can correct verbally and are very often accurate and reliable.

I'm not saying there's been no improvement at all. I personally wouldn't categorize it as staggering, but we can agree to disagree on that. I find the improvements to be uneven in the sense that every time I try a new model I can find use cases where its an improvement over previous versions but I can also find use cases where it feels like a serious regression. Our differences in how we categorize the amount of impr…

Hilarious. Over two years we went from LLMs being slow and not very capable of solving problems to models that are incredibly fast, cheap and able to solve problems in different domains.

Re: GPT-4.5

#815

Funny times. Sonnet 3.7 launches and there is big hype... but complaints start to surface on r/cursor that it is doing too much, is too confident, has no personality. I wonder if 4.5 will be the reverse, an under-hyped launch, but a dawning realisation that it is incredibly useful. Time will tell!

I share the sentiment, as far as I've used it, Sonnet 3.7 is a downgrade and I use Sonnet 3.5 instead. 3.7 tends to overlook critical parts of the query and confidently answers with irrelevant garbage. I'm not sure how QA is done on LLM-s, but I for one definitely feel like the ball was dropped somewhere.

Re: GPT-4.5

#816
post #339

Earlier quoted context omitted.

GPT-4.5 may be an awesome model, some say!

Claude just got a version bump from 3.5 to 3.7. Quite a few people have been asking when OpenAI will get a version bump as well, as GPT 4 has been out "what feels like forever" in the words of a specialist I speak with. Releasing GPT 4.5 might simply be a reaction to Claude 3.7.

since 4o openai has released:

o1 preview. o1 mini. o1. sora. o3-mini <- very good at code

Re: GPT-4.5

#817

Earlier quoted context omitted.

> "Early testing shows that interacting with GPT‑4.5 feels more natural. Its broader knowledge base, improved ability to follow user intent, and greater “EQ” make it useful for tasks like improving writing, programming, and solving practical problems. We also expect it to hallucinate less." "Early testing doesn't show that it hallucinates less, but we expect that putting that sentence nearby will lead you to draw a c…

That's some top-tier sales work right there. I suck at and hate writing the mildly deceptive corporate puffery that seems to be in vogue. I wonder if GPT-4.5 can write that for me or if it's still not as good at it as the expert they paid to put that little gem together.

Good sales lines are like prompt injection for the human mind.

Re: GPT-4.5

#819

It is interesting that they are focusing a large part of this release on the model having a higher "EQ" (Emotional Quotient). We're far from the days of "this is not a person, we do not want to make it addictive" and getting a firm foot on the territory of "here's your new AI friend". This is very visible in the example comparing 4o with 4.5 when the user is complaining about failing a test, where 4o's response is wh…

The whole robotic, monotone, helpful assistant thing was something these companies had to actively hammer in during the post-training stage. It's not really how LLMs will sound by default after pre-training. I guess they're caring less and less about that effort especially since it hurts the model in some ways like creative writing.

Maybe, but I'm not sure how much the style is deliberate vs. a consequence of the post-training tasks like summarization and problem solving. Without seeing the post-training tasks and rating systems it's hard to judge if it's a deliberate style or an emergent consequence of other things.

But it's definitely the case that base models sound more human than instruction-tuned variants. And the shift isn't just vocabulary, it's also in grammar and rhetorical style. There's a shift toward longer words, but also participial phrases, phrasal coordination (with "and" and "or"), and nominalizations (turning adjectives/adverbs into nouns, like "development" or "naturalness"). https://arxiv.org/abs/2410.16107

Re: GPT-4.5

#820

Earlier quoted context omitted.

That $7 trillion dollar ask pushed me from skeptical to full-on eye-roll emoji land— the dude is clearly a narcissist with delusions of grandeur— but it’s getting worse. Considering the $200 pro subscription was significantly unprofitable before this model came out, imagine how astonishingly expensive this model must be to run at many times that price.

Or, the model is nowhere as expensive as in the api pricing and they want to pump the user value of their pro subscription artificially?

Sell an unlimited premium enterprise subscription to every CyberTruck owner, including a huge red ostentatious swastika-shaped back window sticker [but definitely NOT actually an actual swastika, merely a Roman Tetraskelion Strength Symbol] bragging about how much they're spending.
Post reply on HN