Live data from Hacker News

GPT-4.5

openai.com

421–430 of 1001 posts

Re: GPT-4.5

#421

Earlier quoted context omitted.

And LLM's already have tons of productive uses. The biggest ones are probably still waiting, though. But this is about one particular price/performance ratio. You need to build things before you can see how the market responds. You say it's "not good business" but that's entirely wrong. It's excellent business. It's the only way to go about it, in fact. Finding product-market fit is a process. Companies aren't omnisc…

> And LLM's already have tons of productive uses. I disagree strongly with that. Right now they are fun toys to play with, but not useful tools, because they are not reliable. If and when that gets fixed, maybe they will have productive uses. But for right now, not so much.

"it only needs to be good enough" there are tons of productive uses for them. Reliable, much less. But productive? Tons

Re: GPT-4.5

#422

Earlier quoted context omitted.

Now the real question about AI automation starts. Is it cheaper to pay a human to do the task or a AI company?

It still not smart enough to replace for example customer service.

It's absolutely able to replace the majority of customer service volume which is full of mundane questions.

Re: GPT-4.5

#423
post #309
post #280

Earlier quoted context omitted.

It hallucinates at 37% on SimpleQA yeah, which is a set of very difficult questions inviting hallucinations. Claude 3.5 Sonnet (the June 2024 editiom, before October update and before 3.7) hallucinated at 35%. I think this is more of an indication of how behind OpenAI has been in this area.

Are the benchmarks known ahead of time? Could the answer to the benchmarks be in the training data?

They've been caught in the past getting benchmark data under the table, if they got caught once they're probably doing it even more

Re: GPT-4.5

#424

Earlier quoted context omitted.

And LLM's already have tons of productive uses. The biggest ones are probably still waiting, though. But this is about one particular price/performance ratio. You need to build things before you can see how the market responds. You say it's "not good business" but that's entirely wrong. It's excellent business. It's the only way to go about it, in fact. Finding product-market fit is a process. Companies aren't omnisc…

> And LLM's already have tons of productive uses. I disagree strongly with that. Right now they are fun toys to play with, but not useful tools, because they are not reliable. If and when that gets fixed, maybe they will have productive uses. But for right now, not so much.

Hello? Do you have a pulse? LLMs accomplish like 90% of everything I do now so I don’t have to do it…

Explain what this code syntax means…

Explain what this function does…

Write a function to do X…

Respond to my teammates in a Jira ticket explaining why it’s a bad idea to create a repo for every dockerfile…

My teammate responded with X write a rebuttal…

… and the list goes on … like forever

Re: GPT-4.5

#425

Earlier quoted context omitted.

The price really is eye watering. At a glance, my first impression is this is something like Llama 3.1 405B, where the primary value may be realized in generating high quality synthetic data for training rather than direct use. I keep a little google spreadsheet with some charts to help visualize the landscape at a glance in terms of capability/price/throughput, bringing in the various index scores as they become ava…

[flagged]

Nobody comes to HN to read what ChatGPT thinks about something in the comments

Re: GPT-4.5

#426
post #384
post #365

Earlier quoted context omitted.

>"I also agree with researchers like Yann LeCun or François Chollet that deep learning doesn't allow models to generalize properly to out-of-distribution data—and that is precisely what we need to build artificial general intelligence." I think "generalize properly to out-of-distribution data" is too weak of criteria for general intelligence (GI). GI model should be able to get interested about some particular area,…

most humans are generally intelligent but can't do what you just said AGI should do...

Excluding the realtime-iness, humans do at least possess the capacity to do so.

Besides, humans are capable of rigorous logic (which I believe is the most crucial aspect of intelligence) which I don’t think an agent without a proof system can do.

Re: GPT-4.5

#427

GPT-4.5 Preview scored 45% on aider's polyglot coding benchmark [0]. OpenAI describes it as "good at creative tasks" [1], so perhaps it is not primarily intended for coding. 65% Sonnet 3.7, 32k think tokens (SOTA) 60% Sonnet 3.7, no thinking 48% DeepSeek V3 45% GPT 4.5 Preview [0] https://aider.chat/docs/leaderboards/ [1] https://platform.openai.com/docs/models#gpt-4-5

I was waiting for your comment and wow... that's bad. I guess they are ceding the LLMs for coding market to Anthropic? I remember seeing an industry report somewhere and it claimed software development is the largest user of LLMs, so it seems weird to give up in this area.

4.5 lies on a different path than their STEM models.

o3-mini is an extremely powerful coding model and unquestionably is in the same league as 3.7. o3 is still the top stem overall model.

Re: GPT-4.5

#428
post #177

Earlier quoted context omitted.

You're the one buying him the underwear. Don't index funds outperform managed investing? I think especially after accounting for fees, but possibly even after accounting that 50% of money managers are below average.

Depends who's pitch deck you're reading. Warren Buffett didn't get rich waiting on index funds.

And for every Warren Buffet, there are a number of equally competent people who have been less lucky and gone broke taking risks.

Re: GPT-4.5

#429

Earlier quoted context omitted.

I opened your link in a new tab and looked at it a couple minutes later. By then I forgot which was o and which was .5 I honestly couldn't decide which I prefer

I definitely prefer the 4.5, but that might just be because it sounds 'less like ChatGPT', ironically.

It just feels natural to me. The person knows the language but they are not trying to sound smart by using words that might have more impact "based on the words dictionary definition"

GPT 4.5 does feel like it is a step forward in producing natural language, and if they use it to provide reinforcement learning, this might have significant impact in the future smaller models.

Re: GPT-4.5

#430
post #339

Earlier quoted context omitted.

> "Early testing shows that interacting with GPT‑4.5 feels more natural. Its broader knowledge base, improved ability to follow user intent, and greater “EQ” make it useful for tasks like improving writing, programming, and solving practical problems. We also expect it to hallucinate less." "Early testing doesn't show that it hallucinates less, but we expect that putting that sentence nearby will lead you to draw a c…

GPT-4.5 may be an awesome model, some say!

Everybody knows that we're all saying it! That's what I hear from people who should know. And they are so excited about the possibilities!
Post reply on HN