Live data from Hacker News

GPT-4.5

openai.com

481–490 of 1001 posts

Re: GPT-4.5

#481
post #73

Earlier quoted context omitted.

I would like to see a humor test. So far, I have not seen any model response that has made me laugh.

How does the following stand-up routine by Claude 3.7 Sonnet work for you? https://gally.net/temp/20250225claudestandup2.html

I chuckled.

Now you just need a Pro subscription to get Sora generate a video to go along with this and post it to YouTube and rake in the views (and the money that goes along with it).

Re: GPT-4.5

#482
post #216

Earlier quoted context omitted.

To be fAir, SC is trying to do things that no one else done in a context of a single game. I applaud their dedication, but I won't be buying JPGs of a ship for 2k.

Give the same amount of money to a better team and you'd get a better (finished) game. So the allocation of capital is wrong in this case. People shouldn't pre-order stuff. The misallocation of capital also applies to GPT-4.5/OpenAI at this point.

Yeah, I wonder what the Frontier devs could have done with $500M USD. More than $500M USD and 12+ years of development and the game is still in such a sorry state it barely qualifies as little more than a tech demo.

Re: GPT-4.5

#483
I just played with the preview through the API. I asked it to refactor a fairly simple dashboard made with HTML, css and JavaScript.

First time it confused css and JavaScript, then spat out code which broke the dashboard entirely.

Then it charged me $1.53 for the privilege.

Re: GPT-4.5

#484

Earlier quoted context omitted.

Because Russia did undeniably open hostilities? They even admitted to this both times. The second admission being in the form of announcing a “special military operation” when the ceasefire was still active. We also have photographic evidence of them building forces on a border during a ceasefire and then invading. This is like responding to: “did Alexander the Great invade Egypt” by going on a diatribe about how muc…

Okay - but EXACTLY how wrong (or not correct) is the second answer? Please tell me precisely on a 0-1 floating scale, where 0 is "yes" and "no".

One might say, if this were a test being done by a human in a history class, that the answer is 100% incorrect given the actual record of events and failure of statement to mention that actual record. You can argue the causes but that’s not the question.

Re: GPT-4.5

#485
post #339

Earlier quoted context omitted.

> "Early testing shows that interacting with GPT‑4.5 feels more natural. Its broader knowledge base, improved ability to follow user intent, and greater “EQ” make it useful for tasks like improving writing, programming, and solving practical problems. We also expect it to hallucinate less." "Early testing doesn't show that it hallucinates less, but we expect that putting that sentence nearby will lead you to draw a c…

GPT-4.5 may be an awesome model, some say!

Claude just got a version bump from 3.5 to 3.7. Quite a few people have been asking when OpenAI will get a version bump as well, as GPT 4 has been out "what feels like forever" in the words of a specialist I speak with.

Releasing GPT 4.5 might simply be a reaction to Claude 3.7.

Re: GPT-4.5

#486
instead of these random IDs they should label them to make sense for the end user. i have no idea which one to select for what i need. and do they really differ that much by use case?

Re: GPT-4.5

#488
post #342

I got gpt-4.5-preview to summarize this discussion thread so far (at 324 comments): hn-summary.sh 43197872 -m gpt-4.5-preview Using this script: https://til.simonwillison.net/llms/claude-hacker-news-themes... Here's the result: https://gist.github.com/simonw/5e9f5e94ac8840f698c280293d399... It took 25797 input tokens and 1225 input tokens, for a total cost (calculated using https://tools.simonwillison.net/llm-prices…

Maybe it's just confirmation bias but the language in your result output seems higher quality that previous models. Seems more natural and eloquent.

Re: GPT-4.5

#489
post #448

Earlier quoted context omitted.

It’d be great if someone would do that with the same data and prompt to other models. I did like the formatting and attributions but didn’t necessarily want attributions like that for every section. I’m also not sure if it’s fully matching what I’m seeing in the thread but maybe the data I’m seeing is just newer.

Good call. Here's the same exact prompt run against: GPT-4o: https://gist.github.com/simonw/592d651ec61daec66435a6f718c06... GPT-4o Mini: https://gist.github.com/simonw/cc760217623769f0d7e4687332bce... Claude 3.7 Sonnet: https://gist.github.com/simonw/6f11e1974e4d613258b3237380e0e... Claude 3.5 Haiku: https://gist.github.com/simonw/c178f02c97961e225eb615d4b9a1d... Gemini 2.0 Flash: https://gist.github.com/simonw/0c6f…

At a glance, none of these appear to be meaningfully worse than GPT-4.5

Re: GPT-4.5

#490
post #434
post #415

Earlier quoted context omitted.

Words have valence, and valence reflects the state of emotional being of the user. This model appears to understand that better and responds like it’s in a therapeutic conversation and not composing an essay or article. Perhaps they are/were going for stealth therapy-bot with this.

But there is no actual empathy, it isn’t possible.

You think that it isn’t possible to have an emotional model of a human? Why, because you think it is too complex?

Empathy done well seems like 1:1 mapping at an emotional level, but that doesn’t imply to me that it couldn’t be done at a different level of modeling. Empathy can be done poorly, and then it is projecting.

Post reply on HN