Live data from Hacker News

GPT-4.5

openai.com

821–830 of 1001 posts

Re: GPT-4.5

#821
post #390

Earlier quoted context omitted.

I think the trick is observing what is “better” in this model. EQ is supposed to be “better” than 4o, according to the prose. However, how can an LLM have emotional-anything? LLMs are a regurgitation machine, emotion has nothing to do with anything.

Imagine two greeting cards. One says “I’m so sorry for your loss”, and the other says “Everyone dies, they weren’t special”. Does one of these have a higher EQ, despite both being ink and paper and definitely not sentient? Now, imagine they were produced by two different AIs. Does one AI demonstrate higher EQ? The trick is in seeing that “EQ of a text response” is not the same thing as “EQ of a sentient being”

i agree with you. i think it is dishonest for them to post train 4.5 to feign sympathy when someone vents to it. its just weird. they showed it off in the demo.

Re: GPT-4.5

#822

Earlier quoted context omitted.

It's absolutely able to replace the majority of customer service volume which is full of mundane questions.

Such brutal reductionism: how do you calculate an ever growing percentage of customers so pissed at this terrible service that you lose customers forever? Not just one company losing customers... but an entire population completely distrusting and pulling back from any and all companies pulling this trash

Huh? Most call centers these days already use ivr systems and they absolutely are terrible experiences. I along with most people would happily speak with a LLM backed agent to resolve issues.

The CS is already a wreck and LLMs beat an ivr any day of the week and have the ability to offer real triaging ability.

The only people getting upset are the luddites like yourself.

Re: GPT-4.5

#823
post #342

I got gpt-4.5-preview to summarize this discussion thread so far (at 324 comments): hn-summary.sh 43197872 -m gpt-4.5-preview Using this script: https://til.simonwillison.net/llms/claude-hacker-news-themes... Here's the result: https://gist.github.com/simonw/5e9f5e94ac8840f698c280293d399... It took 25797 input tokens and 1225 input tokens, for a total cost (calculated using https://tools.simonwillison.net/llm-prices…

Seems to have trouble recognizing sarcasm:

"For example, there are now a bunch of vendors that sell 'respond to RFP' AI products... paying 30x for marginally better performance makes perfect sense." — hn_throwaway_99 (an uncommon opinion supporting possible niche high-cost uses).

Re: GPT-4.5

#824

Earlier quoted context omitted.

> "Early testing shows that interacting with GPT‑4.5 feels more natural. Its broader knowledge base, improved ability to follow user intent, and greater “EQ” make it useful for tasks like improving writing, programming, and solving practical problems. We also expect it to hallucinate less." "Early testing doesn't show that it hallucinates less, but we expect that putting that sentence nearby will lead you to draw a c…

What is happening to hacker news? I can understand skepticism of new tools like this but the response I see is just so uncurious.

Trough of disillusionment.

A lot of folks here their stock portfolio propped up by AI companies but think they've been overhyped (even if only indirectly through a total stock index). Some were saying all along that this has been a bubble but have been shouted down by true believers hoping for the singularly to usher in techno-utopia.

These signs that perhaps it's been a bit overhyped are validation. The singularly worshipers are much less prominent and so the comments rising to the top are about negatives and not positives.

Ten years from now everyone will just take these tools for granted as much as we take search for granted now.

Re: GPT-4.5

#825

Earlier quoted context omitted.

Please tell me how we objectively determine how correct something is when you ask an LLM: "Was Russia the aggressor in the current Ukraine / Russia conflict?" One LLM says: "Yes." The other says: "Well, it's hard to say because what even is war? And there's been conflict forever, and you have to understand that many people in Russia think there is no such thing as Ukraine and it's always actually just been Russia. Ho…

Because Russia did undeniably open hostilities? They even admitted to this both times. The second admission being in the form of announcing a “special military operation” when the ceasefire was still active. We also have photographic evidence of them building forces on a border during a ceasefire and then invading. This is like responding to: “did Alexander the Great invade Egypt” by going on a diatribe about how muc…

We'll agree to disagree. /s

Re: GPT-4.5

#826

In many ways I'm not an OpenAI fan (but I need to recognize their many merits). At the same time, I believe people are missing what they tried to do with GPT 4.5: it was needed and important to explore the pre-training scaling law in that direction. A gift to science, however selfist it could be.

> A gift to science This is hardly recognizable as science. edit: Sorry, didn't feel this was a controversial opinion. What I meant to say was that for so-called science, this is not reproducible in any way whatsoever. Further, this page in particular has all the hallmarks of _marketing_ copy, not science. Sometimes a failure is just a failure, not necessarily a gift. People could tell scaling wasn't working well bef…

if i understand correctly your argument, then i would say that it is very recognizable as science

>People could tell scaling wasn't working well before the release of GPT 4.5

Yes, on quick glance it seems so from 2020 openai research into scaling laws.

Scaling apparently didn't work well, so the theory about scaling not working well failed to be falsified. It's science.

Re: GPT-4.5

#827

Earlier quoted context omitted.

I don’t mean to disagree too strongly, but just to illustrate another perspective: I don’t feel this is a weak result. Consider if you built a new version that you _thought_ would perform much better, and then you found that it offered marginal-but-not-amazing improvement over the previous version. It’s likely that you will keep iterating. But in the meantime what do you do with your marginal performance gain? Do you…

I've worked for very large software companies, some of the biggest products ever made, and never in 25 years can I recall us shipping an update we didn't know was an improvement. The idea that you'd ship something to hundreds of millions of users and say "maybe better, we're not sure, let us know" is outrageous.

How many times were you in the position to ship something in cutting edge AI? Not trying to be snarky and merely illustrating the point that this is a unique situation. I’d rather they release it and let willing people experiment than not release it at all.

Re: GPT-4.5

#828

Earlier quoted context omitted.

Now the real question about AI automation starts. Is it cheaper to pay a human to do the task or a AI company?

Humans have all sorts of issues you have to deal with. Being hungover, not sleeping well, having a personality, being late to work, not being able to work 24/7, very limited ability to copy them. If there's a soulless generic office-droidGPT that companies could hire that would never talk back and would do all sorts of menial work without needing breaks or to use the bathroom, I don't know that we humans stand a chan…

This is the ultimate business model.

Re: GPT-4.5

#829

Earlier quoted context omitted.

I've worked for very large software companies, some of the biggest products ever made, and never in 25 years can I recall us shipping an update we didn't know was an improvement. The idea that you'd ship something to hundreds of millions of users and say "maybe better, we're not sure, let us know" is outrageous.

Maybe accidental, but I feel you’ve presented a straw man. We’re not discussing something that _may be_ better. It _is_ better. It’s not as big an improvement as previous iterations have been, but it’s still improvement. My claim is that reasonable people might still ship it.

You’re right and... the real issue isn’t the quality of the model or the economics (even when people are willing to pay up). It is the scarcity of GPU compute. This model in particular is sucking up a lot of inference capacity. They are resource constrained and have been wanting more GPUs but they’re only so many going around (demand is insane and keeps growing).
Post reply on HN