Live data from Hacker News

GPT-4.5

openai.com

931–940 of 1001 posts

Re: GPT-4.5

#931

Earlier quoted context omitted.

Is it fair to still call LLMs stochastic parrots now that they are enriched with reasoning? Seems to me that the simple procedure of large-scale sampling + filtering makes it immediately plausible to get something better than the training distribution out of the LLM. In that sense the parrot metaphor seems suddenly wrong. I don’t feel like this binary shift is adequately accounted for among the LLM cynics.

They are not enriched with reasoning, it's just snake oil, I'm afraid.

I'd like to say that with my gut but, at the same time, I've not actually seen a solid definition of what process would define reasoning to say "and this could never be it in any way!". If anything, "a iterative noisy search of similar outputs" now feels at least a big part of what the process of reasoning might need to involve.

Re: GPT-4.5

#932

Earlier quoted context omitted.

Claude just got a version bump from 3.5 to 3.7. Quite a few people have been asking when OpenAI will get a version bump as well, as GPT 4 has been out "what feels like forever" in the words of a specialist I speak with. Releasing GPT 4.5 might simply be a reaction to Claude 3.7.

Feels like when Slackware bumped their Linux version from 4 to 7 just to show they were not falling behind the rest. Wow, I'm old.

Wasn't that the release that they put up the fake IIS page?

Now get off my lawn ))

Re: GPT-4.5

#933

Earlier quoted context omitted.

I begin to believe LLM benchmarks are like european car mileage specs. They say its 4 Liter / 100km but everyone knows it's at least 30% off (same with WLTP for EVs).

Those numbers are not off. They are tested on tracks. You need to remove your shoe and drive with like two toes to get the speed just right, though. Test drivers I have done this with takes off their shoes or use ballerina shoes.

Cruise control?

Re: GPT-4.5

#934

Earlier quoted context omitted.

That $7 trillion dollar ask pushed me from skeptical to full-on eye-roll emoji land— the dude is clearly a narcissist with delusions of grandeur— but it’s getting worse. Considering the $200 pro subscription was significantly unprofitable before this model came out, imagine how astonishingly expensive this model must be to run at many times that price.

Or, the model is nowhere as expensive as in the api pricing and they want to pump the user value of their pro subscription artificially?

Considering that’s the exact opposite of their strategy to date, and they haven’t done anything to indicate that was the case, and they talked about how huge and expensive the model was to run, that is the less reasonable assumption by a mile.

Re: GPT-4.5

#935

Earlier quoted context omitted.

Is it fair to still call LLMs stochastic parrots now that they are enriched with reasoning? Seems to me that the simple procedure of large-scale sampling + filtering makes it immediately plausible to get something better than the training distribution out of the LLM. In that sense the parrot metaphor seems suddenly wrong. I don’t feel like this binary shift is adequately accounted for among the LLM cynics.

it was never fair to call them stochastic parrots and anybody who is paying any attention knows that sequence models can generalize at least partially OOD

Weird the section of people wanting fairness to LLMs.

If it makes you feel better, I'd say the Eliza Effect is good evidence human have a lot of "stochastic parrot" in them also. And there's no reason that being stochastic parrot means something can't generalize.

The thing with these terms is LLMs are distinctly new things. Even blind men looking at elephants can improve their performance with good terminology and by listening to each other. "Effective searchers", "question answers" and "stochastic parrots" are useful term just 'cause the describe concrete behaviors - notably "stochastic parrots" gives some idea of the "no particular goal" quality of LLMs (will happily be NAZIs, pacifists or communists given the proper context). On the other hand, "intelligent" gives no good clues since humans haven't really defined the term for themselves and it is a synonym for good, worthy or capable (giving the machine a prize rather than looking at it).

Re: GPT-4.5

#936

Earlier quoted context omitted.

> "Early testing shows that interacting with GPT‑4.5 feels more natural. Its broader knowledge base, improved ability to follow user intent, and greater “EQ” make it useful for tasks like improving writing, programming, and solving practical problems. We also expect it to hallucinate less." "Early testing doesn't show that it hallucinates less, but we expect that putting that sentence nearby will lead you to draw a c…

What is happening to hacker news? I can understand skepticism of new tools like this but the response I see is just so uncurious.

Just like cryptocurrency. For a brief moment, HN worshiped at the altar of the blockchain. This technology was going to revolutionize the world and democratize everything. Then some negative financial stuff happened, and people realized that most of cryptocurrency is puffery and scams. Now you can hardly find a positive comment on cryptocurrency.

Re: GPT-4.5

#937
post #609
post #448

Earlier quoted context omitted.

Good call. Here's the same exact prompt run against: GPT-4o: https://gist.github.com/simonw/592d651ec61daec66435a6f718c06... GPT-4o Mini: https://gist.github.com/simonw/cc760217623769f0d7e4687332bce... Claude 3.7 Sonnet: https://gist.github.com/simonw/6f11e1974e4d613258b3237380e0e... Claude 3.5 Haiku: https://gist.github.com/simonw/c178f02c97961e225eb615d4b9a1d... Gemini 2.0 Flash: https://gist.github.com/simonw/0c6f…

I actually think the Claude 3.7 Sonnet summary is better.

yeah I liked it too, especially for 10x less the price lol

Re: GPT-4.5

#938

Earlier quoted context omitted.

AI in general is increasingly a solution in search of a problem, so this seems about right.

Only in the same sense as electricity is. The main tools apply to almost any activity humans do. It's already obvious that it's the solution to X for almost any X, but the devil is in the details - i.e. picking specific, simplest problems to start with.

No, in the sense that blockchain is. This is just the latest in a long history of tech fads propelled by wishful thinking and unqualified grifters.

It is the solution to almost nothing, but is being shoehorned into every imaginable role by people who are blind to its shortcomings, often wilfully. The only thing that's obvious to me is that a great number of people are apparently desperate for a tool to do their thinking for them, no matter how garbage the result is. It's disheartening to realize that so many people consider using their own brain to be such an intolerable burden.

Re: GPT-4.5

#939

Earlier quoted context omitted.

As someone who is terrified of agentic ASI, I desperately hope this is true. We need more time to figure out alignment.

"alignment" is a bs term made up to deflect blame from the overpromises the AI companies made to hype up their product to obtain their valuations.

Big take given how much AI companies hate alignment folks.

Re: GPT-4.5

#940

Earlier quoted context omitted.

In the second handpicked example they give, GPT-4.5 says that "The Trojan Women Setting Fire to Their Fleet" by the French painter Claude Lorrain is renowned for its luminous depiction of fire. That is a hallucination. There is no fire at all in the painting, only some smoke. https://en.wikipedia.org/wiki/The_Trojan_Women_Set_Fire_to_t...

AI crash is gonna lead to decade long winter

On the bright side, at least we'll be able to warm our hands by the waste heat of the GPUs.
Post reply on HN