Live data from Hacker News

GPT-4.5

openai.com

841–850 of 1001 posts

Re: GPT-4.5

#841

Is it official then? Most of us have been waiting for this moment for a while. The transformer architecture as it is currently understood can't be milked any further. Many of us knew this since last year. GPT-5 delays eventually led to non-tech voices to suggest likewise. But we all held our final decision until the next big release from OpenAI as Sam Altman has been making claims about AGI entering the workforce thi…

I'm not convinced that LLMs in their current state are really making anyone's lives much better though. We really need more research applications for this technology for that to become apparent. Polluting the internet with regurgitated garbage produced by a chat bot does not benefit the world. Increasing the productivity of software developers does not help to the world. Solving more important problems should be the priority for this type of AI research & development.

Re: GPT-4.5

#842

Is it official then? Most of us have been waiting for this moment for a while. The transformer architecture as it is currently understood can't be milked any further. Many of us knew this since last year. GPT-5 delays eventually led to non-tech voices to suggest likewise. But we all held our final decision until the next big release from OpenAI as Sam Altman has been making claims about AGI entering the workforce thi…

As someone who is terrified of agentic ASI, I desperately hope this is true. We need more time to figure out alignment.

Re: GPT-4.5

#843
I've been using 4.5 for the better part of the day.

I also have access to o3-mini-high and o1-pro.

I don't get it. For general purposes and for writing, 4.5 is no better than o3-mini. It may even be worse.

I'd go so far as to say that Deepseek is actually better than 4.5 for most general purpose use cases.

I seriously don't understand what they're trying to achieve with this release.

Re: GPT-4.5

#845

Earlier quoted context omitted.

Have you tried copying the compilation errors back into the prompt? In my experience eventually the result is correct. If not then I shrink the surface area that the model is touching and try again.

yes ofcourse. it then proceeds to agree that what it told me was indeed stupid and proceeds to give me something even worse. I would love to see a video of ppl using this in real projects ( even if its open source) . I am tried of ppl claiming moon and stars after trying it on toy projects.

Yeah that's what happens. It can recreate anything it's been trained on - which is a lot - but you'll definitely fall into these "Oh, I see the issue now" loops when doing anything not in the training set.

Re: GPT-4.5

#846

I've been using 4.5 for the better part of the day. I also have access to o3-mini-high and o1-pro. I don't get it. For general purposes and for writing, 4.5 is no better than o3-mini. It may even be worse. I'd go so far as to say that Deepseek is actually better than 4.5 for most general purpose use cases. I seriously don't understand what they're trying to achieve with this release.

this model does have a niche use-case: since its so large it does have a lot more knowledge and hallucinates much less. for example as a test question I asked it to list the best restaurants in my small town. and all of them existed. none of the other llms get this right.

Re: GPT-4.5

#847
post #841

Is it official then? Most of us have been waiting for this moment for a while. The transformer architecture as it is currently understood can't be milked any further. Many of us knew this since last year. GPT-5 delays eventually led to non-tech voices to suggest likewise. But we all held our final decision until the next big release from OpenAI as Sam Altman has been making claims about AGI entering the workforce thi…

I'm not convinced that LLMs in their current state are really making anyone's lives much better though. We really need more research applications for this technology for that to become apparent. Polluting the internet with regurgitated garbage produced by a chat bot does not benefit the world. Increasing the productivity of software developers does not help to the world. Solving more important problems should be the…

> Solving more important problems should be the priority for this type of AI research & development.

Which problem spaces do you think are underserved in this aspect?

Re: GPT-4.5

#848

Earlier quoted context omitted.

The link has data. The link shows a significant reduction. grep hallucination, or, https://imgur.com/a/mkDxe78 .

I really doubt LLM benchmarks are reflective of real world user experience ever since they claimed GPT-4o hallucinated less than the original GPT-4.

I don't have an accurate benchmark, but in my personal experience, gpt4o hallucinates substantially less than gpt4. We solved a ton of hallucination issues just by upgrading to it...

Re: GPT-4.5

#849
post #846

I've been using 4.5 for the better part of the day. I also have access to o3-mini-high and o1-pro. I don't get it. For general purposes and for writing, 4.5 is no better than o3-mini. It may even be worse. I'd go so far as to say that Deepseek is actually better than 4.5 for most general purpose use cases. I seriously don't understand what they're trying to achieve with this release.

this model does have a niche use-case: since its so large it does have a lot more knowledge and hallucinates much less. for example as a test question I asked it to list the best restaurants in my small town. and all of them existed. none of the other llms get this right.

I tried the same thing with companies in my industry ("list active companies in the field of X") and it came back with a few that have been shuttered for years, in one case for nearly two decades.

I'm really not seeing better performance than with o3-mini.

If anything, the new results ("list active companies in the field of X") are actually worse than what I'd get with o3-mini, because the 4.5 response is basically the post-SEO Google first page (it appears to default to mentioning the companies that rank most highly on Google,) whereas the o3 response was more insightful and well-reasoned.

Re: GPT-4.5

#850

Earlier quoted context omitted.

Imagine two greeting cards. One says “I’m so sorry for your loss”, and the other says “Everyone dies, they weren’t special”. Does one of these have a higher EQ, despite both being ink and paper and definitely not sentient? Now, imagine they were produced by two different AIs. Does one AI demonstrate higher EQ? The trick is in seeing that “EQ of a text response” is not the same thing as “EQ of a sentient being”

i agree with you. i think it is dishonest for them to post train 4.5 to feign sympathy when someone vents to it. its just weird. they showed it off in the demo.

Why? The choice to not do the post training would be every bit as intentional, and no different than post training to make it less sympathetic.

This is a designed system. The designers make choices. I don’t see how failing to plan and design for a common use case would be better.

Post reply on HN