Live data from Hacker News

GPT-4.5

openai.com

881–890 of 1001 posts

Re: GPT-4.5

#881

Is it official then? Most of us have been waiting for this moment for a while. The transformer architecture as it is currently understood can't be milked any further. Many of us knew this since last year. GPT-5 delays eventually led to non-tech voices to suggest likewise. But we all held our final decision until the next big release from OpenAI as Sam Altman has been making claims about AGI entering the workforce thi…

It's worth pointing out that GPT-4.5 seems focused on better pre-training and doesn't include reasoning. I think GPT-5 - if/when it happens - will be 4.5 with reasoning, and as such it will feel very different. The barrier, is the computational cost of it. Once 4.5 gets down to similar costs to 4.0 - which could be achieved through various optimization steps (what happened to the ternary stuff that was published last…

Is it fair to still call LLMs stochastic parrots now that they are enriched with reasoning? Seems to me that the simple procedure of large-scale sampling + filtering makes it immediately plausible to get something better than the training distribution out of the LLM. In that sense the parrot metaphor seems suddenly wrong.

I don’t feel like this binary shift is adequately accounted for among the LLM cynics.

Re: GPT-4.5

#882

Is it official then? Most of us have been waiting for this moment for a while. The transformer architecture as it is currently understood can't be milked any further. Many of us knew this since last year. GPT-5 delays eventually led to non-tech voices to suggest likewise. But we all held our final decision until the next big release from OpenAI as Sam Altman has been making claims about AGI entering the workforce thi…

> Not much consolation to the world's super rich who will lose tons of money once the LLM industry (let us remember that AI is not LLM) falls.

They knew the deal:

“it would be wise to view any investment in OpenAI Global, LLC in the spirit of a donation” and “it may be difficult to know what role money will play in a post-[artificial general intelligence] world.”

Re: GPT-4.5

#883

Earlier quoted context omitted.

In the second handpicked example they give, GPT-4.5 says that "The Trojan Women Setting Fire to Their Fleet" by the French painter Claude Lorrain is renowned for its luminous depiction of fire. That is a hallucination. There is no fire at all in the painting, only some smoke. https://en.wikipedia.org/wiki/The_Trojan_Women_Set_Fire_to_t...

AI crash is gonna lead to decade long winter

There have always been cycles of hype and correction.

I don't see AI going any differently. Some companies will figure out where and how models should be utilized, they'll see some benefit. (IMO, the answer will be smaller local models tailored to specific domains)

Others will go bust. Same as it always was.

Re: GPT-4.5

#884

Earlier quoted context omitted.

In the second handpicked example they give, GPT-4.5 says that "The Trojan Women Setting Fire to Their Fleet" by the French painter Claude Lorrain is renowned for its luminous depiction of fire. That is a hallucination. There is no fire at all in the painting, only some smoke. https://en.wikipedia.org/wiki/The_Trojan_Women_Set_Fire_to_t...

AI crash is gonna lead to decade long winter

It will be upheld as prime example that a whole market can self-hypnotize and ruin the society its based upon out of existence against all future pundits of this very economic system.

Re: GPT-4.5

#885
post #857

Earlier quoted context omitted.

I'm not sure this will ever be solved. It requires both a technical solution and social consensus. I don't see consensus on "alignment" happening any time soon. I think it'll boil down to "aligned with the goals of the nation-state", and lots of nation states have incompatible goals.

I agree unfortunately. I might be a bit of an extremist on this issue. I genuinely think that building agentic ASI is suicidally stupid and we just shouldn’t do it. All the utopian visions we hear from the optimists describe unstable outcomes. A world populated by super-intelligent agents will be incredibly dangerous even if it appears initially to have gone well. We’ll have built a paradise in which we can never rel…

What's the difference between your "agentic AIs" and, say, "script kiddies" or "expert anarchist/black-hat hackers"?

It's been obvious for a while that the narrow-waist APIs between things matter, and apparent that agentic AI is leaning into adaptive API consumption, but I don't see how that gives the agentic client some super-power we don't already need to defend against since before AGI we already have HGI (human general intelligence) motivated to "do bad things" to/through those APIs, both self-interested and nation-state sponsored.

We're seeing more corporate investment in this interplay, trending us towards Snow Crash, but "all you have to do" is have some "I" in API be "dual key human in the loop" to enable a scenario where AGI/HGI "presses the red button" in the oval office, nuclear war still doesn't happen, WarGames or Crimson Tide style.

I'm not saying dual key is the answer to everything, I'm saying, defenses against adversaries already matter, and will continue to. We have developed concepts like air gaps or modality changes, and need more, but thinking in terms of interfaces (APIs) in the general rather than the literal gives a rich territory for guardrails and safeguards.

Re: GPT-4.5

#886
post #788
post #591

Earlier quoted context omitted.

It's nice to know the new Turing test is generating effective VC pitch decks.

Joke's on us, the VC's are using LLM's to evaluate the pitch decks.

Chat-GPT generate a prompt injection attack, embedded in a background image.

Re: GPT-4.5

#887

Earlier quoted context omitted.

there are literally hundreds of extensions and sites that do this the problem is that they are competing each other into the ground hence they go unmaintained very quickly getrecall.ai has been the most mature so far

Hey, check this one out with all the different flavors that existed out there. I think I made something better. https://cofyt.app As far as I am aware, feel free to test it head-to-head. This is better than gecall, and you can chat with a transcript for detailed answers to your prompts

I tried it out, looks nice and clean.

But as I mentioned, my main concern is what will happen in 6 months when you fail to get traction and abandon it. Because that's what happened to previous 5 products I tried which were all "good enough" .

Getrecall seems to have a big enough user base that will actually stick around.

Re: GPT-4.5

#888
post #167

First impression of GPT-4.5: 1. It is very very slow, for some applications where you want real time interactions is just not viable, the text attached below took 7s to generate with 4o, but 46s with GPT4.5 2. The style it writes is way better: it keeps the tone you ask and makes better improvements on the flow. One of my biggest complaints with 4o is that you want for your content to be more casual and accessible bu…

Similar reaction here. I will also note that it seems to know a lot more about me than previous models. I’m not sure if this is a broader web crawl, more space in the model, or more summarization of our chats or a combination, but I asked it to psychoanalyze a problem I’m having in the style of Jacques lacan and it was genuinely helpful and interesting, no interview required first; it just went right at me.

To borrow an iain banks word, the “fragre” def feels improved to me. I think I will prefer it to o1 pro, although I haven’t really hammered on it yet.

Re: GPT-4.5

#889
post #848

Earlier quoted context omitted.

I really doubt LLM benchmarks are reflective of real world user experience ever since they claimed GPT-4o hallucinated less than the original GPT-4.

I don't have an accurate benchmark, but in my personal experience, gpt4o hallucinates substantially less than gpt4. We solved a ton of hallucination issues just by upgrading to it...

How much did you use the original GPT-4-0314?

(And even that was a downgrade compared to the more uncensored pre-release versions, which were comparable to GPT-4.5, at least judging by the unicorn test)

Post reply on HN