Earlier quoted context omitted.
And LLM's already have tons of productive uses. The biggest ones are probably still waiting, though. But this is about one particular price/performance ratio. You need to build things before you can see how the market responds. You say it's "not good business" but that's entirely wrong. It's excellent business. It's the only way to go about it, in fact. Finding product-market fit is a process. Companies aren't omnisc…
> And LLM's already have tons of productive uses. I disagree strongly with that. Right now they are fun toys to play with, but not useful tools, because they are not reliable. If and when that gets fixed, maybe they will have productive uses. But for right now, not so much.
GPT-4.5
421–430 of 1001 posts
Re: GPT-4.5
#422Earlier quoted context omitted.
Now the real question about AI automation starts. Is it cheaper to pay a human to do the task or a AI company?
It still not smart enough to replace for example customer service.
Re: GPT-4.5
#423Earlier quoted context omitted.
It hallucinates at 37% on SimpleQA yeah, which is a set of very difficult questions inviting hallucinations. Claude 3.5 Sonnet (the June 2024 editiom, before October update and before 3.7) hallucinated at 35%. I think this is more of an indication of how behind OpenAI has been in this area.
Are the benchmarks known ahead of time? Could the answer to the benchmarks be in the training data?
Re: GPT-4.5
#424Earlier quoted context omitted.
And LLM's already have tons of productive uses. The biggest ones are probably still waiting, though. But this is about one particular price/performance ratio. You need to build things before you can see how the market responds. You say it's "not good business" but that's entirely wrong. It's excellent business. It's the only way to go about it, in fact. Finding product-market fit is a process. Companies aren't omnisc…
> And LLM's already have tons of productive uses. I disagree strongly with that. Right now they are fun toys to play with, but not useful tools, because they are not reliable. If and when that gets fixed, maybe they will have productive uses. But for right now, not so much.
Explain what this code syntax means…
Explain what this function does…
Write a function to do X…
Respond to my teammates in a Jira ticket explaining why it’s a bad idea to create a repo for every dockerfile…
My teammate responded with X write a rebuttal…
… and the list goes on … like forever
Re: GPT-4.5
#425Earlier quoted context omitted.
The price really is eye watering. At a glance, my first impression is this is something like Llama 3.1 405B, where the primary value may be realized in generating high quality synthetic data for training rather than direct use. I keep a little google spreadsheet with some charts to help visualize the landscape at a glance in terms of capability/price/throughput, bringing in the various index scores as they become ava…
[flagged]
Re: GPT-4.5
#426Earlier quoted context omitted.
>"I also agree with researchers like Yann LeCun or François Chollet that deep learning doesn't allow models to generalize properly to out-of-distribution data—and that is precisely what we need to build artificial general intelligence." I think "generalize properly to out-of-distribution data" is too weak of criteria for general intelligence (GI). GI model should be able to get interested about some particular area,…
most humans are generally intelligent but can't do what you just said AGI should do...
Besides, humans are capable of rigorous logic (which I believe is the most crucial aspect of intelligence) which I don’t think an agent without a proof system can do.
Re: GPT-4.5
#427GPT-4.5 Preview scored 45% on aider's polyglot coding benchmark [0]. OpenAI describes it as "good at creative tasks" [1], so perhaps it is not primarily intended for coding. 65% Sonnet 3.7, 32k think tokens (SOTA) 60% Sonnet 3.7, no thinking 48% DeepSeek V3 45% GPT 4.5 Preview [0] https://aider.chat/docs/leaderboards/ [1] https://platform.openai.com/docs/models#gpt-4-5
I was waiting for your comment and wow... that's bad. I guess they are ceding the LLMs for coding market to Anthropic? I remember seeing an industry report somewhere and it claimed software development is the largest user of LLMs, so it seems weird to give up in this area.
o3-mini is an extremely powerful coding model and unquestionably is in the same league as 3.7. o3 is still the top stem overall model.
Re: GPT-4.5
#428Earlier quoted context omitted.
You're the one buying him the underwear. Don't index funds outperform managed investing? I think especially after accounting for fees, but possibly even after accounting that 50% of money managers are below average.
Depends who's pitch deck you're reading. Warren Buffett didn't get rich waiting on index funds.
Re: GPT-4.5
#429Earlier quoted context omitted.
I opened your link in a new tab and looked at it a couple minutes later. By then I forgot which was o and which was .5 I honestly couldn't decide which I prefer
I definitely prefer the 4.5, but that might just be because it sounds 'less like ChatGPT', ironically.
GPT 4.5 does feel like it is a step forward in producing natural language, and if they use it to provide reinforcement learning, this might have significant impact in the future smaller models.
Re: GPT-4.5
#430Earlier quoted context omitted.
> "Early testing shows that interacting with GPT‑4.5 feels more natural. Its broader knowledge base, improved ability to follow user intent, and greater “EQ” make it useful for tasks like improving writing, programming, and solving practical problems. We also expect it to hallucinate less." "Early testing doesn't show that it hallucinates less, but we expect that putting that sentence nearby will lead you to draw a c…
GPT-4.5 may be an awesome model, some say!