Live data from Hacker News

GPT-4.5

openai.com

581–590 of 1001 posts

Re: GPT-4.5

#581

Earlier quoted context omitted.

I paid for pro to try `o1-pro` and I can't seem to find any use case to justify the insane inference time. `o3-mini-high` seems to do just as well in seconds vs. minutes.

What are you doing with it? For me deep research tasks are where 5 minutes is fine, or something really hard that would take me way more time myself.

I usually throw a lot of context at it and have it write unit tests in a certain style or implement something (with tests) according to a spec.

But the o3-mini-high results have been just as good.

I am fine with Deep Research taking 5-8 minutes, those are usually "reports" I can read whenever.

Re: GPT-4.5

#582

Earlier quoted context omitted.

> We look forward to learning more about its strengths, capabilities, and potential applications in real-world settings. If GPT‑4.5 delivers unique value for your use case, your feedback (opens in a new window) will play an important role in guiding our decision. "We don't really know what this is good for, but spent a lot of money and time making it and are under intense pressure to announce new things right now. If…

> "Early testing shows that interacting with GPT‑4.5 feels more natural. Its broader knowledge base, improved ability to follow user intent, and greater “EQ” make it useful for tasks like improving writing, programming, and solving practical problems. We also expect it to hallucinate less." "Early testing doesn't show that it hallucinates less, but we expect that putting that sentence nearby will lead you to draw a c…

So they made Claude that knows a bit more.

Re: GPT-4.5

#583
post #423

Earlier quoted context omitted.

They've been caught in the past getting benchmark data under the table, if they got caught once they're probably doing it even more

No, they haven't.

They actually have [0]. They were revealed to have had access to the (majority of the) frontierMath problemset while everybody thought the problemset was confidential, and published benchmarks for their o3 models on the presumption that they didn't. I mean one is free to trust their "verbal agreement" that they did not train their models on that, but access they did have and it was not revealed until much later.

[0] https://the-decoder.com/openai-quietly-funded-independent-ma...

Re: GPT-4.5

#584

Earlier quoted context omitted.

> "Early testing shows that interacting with GPT‑4.5 feels more natural. Its broader knowledge base, improved ability to follow user intent, and greater “EQ” make it useful for tasks like improving writing, programming, and solving practical problems. We also expect it to hallucinate less." "Early testing doesn't show that it hallucinates less, but we expect that putting that sentence nearby will lead you to draw a c…

[flagged]

I suspect people downvote you because the tone of your reply makes it seem like you are personally offended and are now firing back with equally unfounded attacks like a straight up "you are lying".

I read the article but can't find the numbers you are referencing. Maybe there's some paper linked I should be looking at? The only numbers I see are from the SimpleQA chart, which are 37.1% vs 61.8% hallucination rate. That's nice but considering the price increase, is it really that impressive? Also, an often repeated criticism is that relying on known benchmarks is "gaming the numbers" and that the real world hallucination rate could very well be higher.

Lastly, the themselves say: > We also expect it to hallucinate less.

That's a fairly neutral statement for a press release. If they were convinced that the reduced hallucination rate is the killer feature that sets this model apart from the competition, they surely would have emphasized that more?

All in all I can understand why people would react with some mocking replies to this.

Re: GPT-4.5

#585
post #452
post #342

I got gpt-4.5-preview to summarize this discussion thread so far (at 324 comments): hn-summary.sh 43197872 -m gpt-4.5-preview Using this script: https://til.simonwillison.net/llms/claude-hacker-news-themes... Here's the result: https://gist.github.com/simonw/5e9f5e94ac8840f698c280293d399... It took 25797 input tokens and 1225 input tokens, for a total cost (calculated using https://tools.simonwillison.net/llm-prices…

Didn't seem to realize that "Still more coherent than the OpenAI lineup" wouldn't make sense out of context. (The actual comment quoted there is responding to someone who says they'd name their models Foo, Bar, Baz.)

Wonder if there’s some pro-OpenAI system prompt getting in the way of that.

Re: GPT-4.5

#586

Earlier quoted context omitted.

ChatGPT had its initial public release November 30th, 2022. That's 820 days to today. The Apple II was first sold June 10, 1977, and Visicalc was first sold October 17, 1979, which is 859 days. So we're right about the same distance in time- the exact equal duration will be April 7th of this year. Going back to the very first commercially available microcomputer, the Altair 8800 (which is not a great match, since tha…

So it’s barely been 2 years. And we’ve already seen pretty crazy progress in that time. Let’s see what a few more years brings.

what crazy progress? how much do you spend on tokens every month to witness the crazy progress that I'm not seeing? I feel like I'm taking crazy pills. The progress is linear at best

Re: GPT-4.5

#587

Earlier quoted context omitted.

ChatGPT had its initial public release November 30th, 2022. That's 820 days to today. The Apple II was first sold June 10, 1977, and Visicalc was first sold October 17, 1979, which is 859 days. So we're right about the same distance in time- the exact equal duration will be April 7th of this year. Going back to the very first commercially available microcomputer, the Altair 8800 (which is not a great match, since tha…

> Visicalc was first sold October 17, 1979, which is 859 days. And it still can't answer simple English-language questions.

it could do math reliably!

Re: GPT-4.5

#589
post #95
post #24

Considering both this blog post and the livestream demos, I am underwhelmed. Having just finished the stream, I had a real "was that all" moment, which on one hand shows how spoiled I've gotten by new models impressing me, but on another feels like OpenAI really struggles to stay ahead of their competitors. What has been shown feels like it could be achieved using a custom system prompt on older versions of OpenAIs m…

rethinking your comment "was that all" I am listening to the stream now and had a thought. Most of the new models that have come out in the past few weeks have been great at coding and logical reasoning. But 4o has been better at creative writing. I am wondering if 4.5 is going to be even better at creative writing than 4o.

if you generate "creative" writing, please tell your audience that it is generated, before asking them to read it.

I do not understand what possible motivation there could be for generating "creative writing" unless you enjoy reading meaningless stories yourself, in which case, be my guest.

Re: GPT-4.5

#590
post #167

First impression of GPT-4.5: 1. It is very very slow, for some applications where you want real time interactions is just not viable, the text attached below took 7s to generate with 4o, but 46s with GPT4.5 2. The style it writes is way better: it keeps the tone you ask and makes better improvements on the flow. One of my biggest complaints with 4o is that you want for your content to be more casual and accessible bu…

I'm wondering if generative AI will ultimately result in a very dense / bullet form style of writing. What we are doing now is effectively this:

bullet_points' = compress(expand(bullet_points))

We are impressed by lots of text so must expand via LLM in order to impress the reader. Since the reader doesn't have time or interest to read the content they must compress it back into bullet points / quick summary. Really, the original bullet points plus a bit more thinking would likely be a better form of communication.

Post reply on HN