Live data from Hacker News

GPT-4.5

openai.com

521–530 of 1001 posts

Re: GPT-4.5

#521

It significantly improves upon GPT-4o on my Extended NYT Connections Benchmark. 22.4 -> 33.7 ( https://github.com/lechmazur/nyt-connections ).

Honest question for you: are these puzzles actually a good way to test the models?

The answers are certainly in the training set, likely many times over.

I’d be curious to see performance on Bracket City, which was featured here on HN yesterday.

Re: GPT-4.5

#522
post #311

Earlier quoted context omitted.

> 1. It is very very slow, ... below took 7s to generate with 4o, but 46s with GPT4.5 This is positively luxurious by o1-pro standards which I'd say average 5 minutes. That said I totally agree even ~45s isn't viable for real-time interactions. I'm sure it'll be optimized. Of course, my comparing it to the highest-end CoT model in [publicly-known] existence isn't entirely fair since they're sort of apples and oranges…

I paid for pro to try `o1-pro` and I can't seem to find any use case to justify the insane inference time. `o3-mini-high` seems to do just as well in seconds vs. minutes.

[deleted]

Re: GPT-4.5

#523
post #280

Earlier quoted context omitted.

According to a graph they provide, it does hallucinate significantly less on at least one benchmark.

It hallucinates at 37% on SimpleQA yeah, which is a set of very difficult questions inviting hallucinations. Claude 3.5 Sonnet (the June 2024 editiom, before October update and before 3.7) hallucinated at 35%. I think this is more of an indication of how behind OpenAI has been in this area.

Benchmarks are not real so 2% is meaningless.

Re: GPT-4.5

#524
post #24

Considering both this blog post and the livestream demos, I am underwhelmed. Having just finished the stream, I had a real "was that all" moment, which on one hand shows how spoiled I've gotten by new models impressing me, but on another feels like OpenAI really struggles to stay ahead of their competitors. What has been shown feels like it could be achieved using a custom system prompt on older versions of OpenAIs m…

I have no idea how they justify $200/month for pro

Re: GPT-4.5

#525
this seems to be a very weak response to sonnet 3.7

- more expensive. alot more expensive

- not a lot of increment improvement

Re: GPT-4.5

#526
post #73

Earlier quoted context omitted.

I would like to see a humor test. So far, I have not seen any model response that has made me laugh.

How does the following stand-up routine by Claude 3.7 Sonnet work for you? https://gally.net/temp/20250225claudestandup2.html

reddit tier humor, truly

it's just regurgitating overly emphasized cliches in a disgustingly enthusiastic tone

Re: GPT-4.5

#527

Earlier quoted context omitted.

Eh, I think o1-pro is by far the most capable model available right now in terms of pure problem solving.

I think Claude has consistently been ahead for a year ish now and is back ahead again for my use cases with 3.7.

[deleted]

Re: GPT-4.5

#528
post #342

I got gpt-4.5-preview to summarize this discussion thread so far (at 324 comments): hn-summary.sh 43197872 -m gpt-4.5-preview Using this script: https://til.simonwillison.net/llms/claude-hacker-news-themes... Here's the result: https://gist.github.com/simonw/5e9f5e94ac8840f698c280293d399... It took 25797 input tokens and 1225 input tokens, for a total cost (calculated using https://tools.simonwillison.net/llm-prices…

Huh. Disregarding the 4.5-specific bit here, a browser extension or possibly website that did this in general could be really useful.

Maybe even something that just noticed whenever you visited a site that had had significant HN discussion in the past, then let you trigger a summary.

Re: GPT-4.5

#529
post #167

First impression of GPT-4.5: 1. It is very very slow, for some applications where you want real time interactions is just not viable, the text attached below took 7s to generate with 4o, but 46s with GPT4.5 2. The style it writes is way better: it keeps the tone you ask and makes better improvements on the flow. One of my biggest complaints with 4o is that you want for your content to be more casual and accessible bu…

What’s the deal with Imgur taking ages to load? Anyone else have this issue in Australia? I just get the grey background with no content loaded for 10+ seconds every time I visit that bloated website.

Ok for me here in aus

Re: GPT-4.5

#530
post #523
post #280

Earlier quoted context omitted.

It hallucinates at 37% on SimpleQA yeah, which is a set of very difficult questions inviting hallucinations. Claude 3.5 Sonnet (the June 2024 editiom, before October update and before 3.7) hallucinated at 35%. I think this is more of an indication of how behind OpenAI has been in this area.

Benchmarks are not real so 2% is meaningless.

Of course not. The point is that the cost difference between the two things being compared is huge, right? Same performance, but not the same cost.
Post reply on HN