Live data from Hacker News

GPT-4.5

openai.com

401–410 of 1001 posts

Re: GPT-4.5

#401
post #144

Earlier quoted context omitted.

> "We don't really know what this is good for, but spent a lot of money and time making it and are under intense pressure to announce new things right now. If you can figure something out, we need you to help us." Where is this quote from?

The quotation marks in the grandparent comment are scare (sneer) quotes and not actual quotation. https://en.m.wikipedia.org/wiki/Scare_quotes > Whether quotation marks are considered scare quotes depends on context because scare quotes are not visually different from actual quotations.

That's not a scare quote. It's just a proposed subtext of the quote. Sarcastic, sure, but no a scare quote, which is a specific kind of thing. (from your linked wikipedia: "... around a word or phrase to signal that they are using it in an ironic, referential, or otherwise non-standard sense.")

Re: GPT-4.5

#403
post #263

Earlier quoted context omitted.

Maybe if they build a few more data centers, they'll be able to construct their machine god. Just a few more dedicated power plants, a lake or two, a few hundred billion more and they'll crack this thing wide open. And maybe Tesla is going to deliver truly full self driving tech any day now. And Star Citizen will prove to have been worth it along along, and Bitcoin will rain from the heavens. It's very difficult to r…

You have it all wrong. The end game is a scalable, reliable AI work force capable of finishing Star Citizen. At least this is the benchmark for super-human general intelligence that I propose.

Man I can't believe that fucking game is still alive and kicking. Tell me they're making good progress, sho_hn

Re: GPT-4.5

#404

Earlier quoted context omitted.

The price really is eye watering. At a glance, my first impression is this is something like Llama 3.1 405B, where the primary value may be realized in generating high quality synthetic data for training rather than direct use. I keep a little google spreadsheet with some charts to help visualize the landscape at a glance in terms of capability/price/throughput, bringing in the various index scores as they become ava…

[flagged]

Don't do this.

Re: GPT-4.5

#405
post #81
post #24

Considering both this blog post and the livestream demos, I am underwhelmed. Having just finished the stream, I had a real "was that all" moment, which on one hand shows how spoiled I've gotten by new models impressing me, but on another feels like OpenAI really struggles to stay ahead of their competitors. What has been shown feels like it could be achieved using a custom system prompt on older versions of OpenAIs m…

> How could they justify that asking price? They're still selling $1 for <$1. Like personal food delivery before it, consumers will eventually need to wake up to this fact - these things will get expensive, fast.

I read this more as "we are releasing a model checkpoint that we didn't optimize yet because Anthropic cranked up the pressure"

Re: GPT-4.5

#406

I cancelled my ChatGPT subscription today in favor of using Grok. It’s literally the difference between me never using ChatGPT to using Grok all the time, and the only way I can explain it is twofold: 1. The output from Grok doesn’t feel constrained. I don’t know how much of this is the marketing pitch of it “not being woke”, but I feel it in its answers. It never tells me it’s not going to return a result or sugarco…

I found Grok's reasoning pretty wack.

I asked it - "Draft a Minnesota Motion in Limine to exclude ..."

It then starts thinking ... User wants a Missouri Motion in Limine ....

Re: GPT-4.5

#408
post #177

Earlier quoted context omitted.

You're the one buying him the underwear. Don't index funds outperform managed investing? I think especially after accounting for fees, but possibly even after accounting that 50% of money managers are below average.

He earns his undies. My returns are almost always modestly above index fund returns after his fees, though like last quarter, he’s very upfront when they’re not. He has good advice for pulling back when things are uncertain. I’m happy to delegate that to him.

[deleted]

Re: GPT-4.5

#409

Earlier quoted context omitted.

I agree it's a hard problem. I think there are a number of tests out there however that are able to objectively test capability and truthfulness. I've read reports that some of the changes that are preferred by human evaluators actually hurt the performance on the more objective tests.

Please tell me how we objectively determine how correct something is when you ask an LLM: "Was Russia the aggressor in the current Ukraine / Russia conflict?" One LLM says: "Yes." The other says: "Well, it's hard to say because what even is war? And there's been conflict forever, and you have to understand that many people in Russia think there is no such thing as Ukraine and it's always actually just been Russia. Ho…

Because Russia did undeniably open hostilities? They even admitted to this both times. The second admission being in the form of announcing a “special military operation” when the ceasefire was still active. We also have photographic evidence of them building forces on a border during a ceasefire and then invading. This is like responding to: “did Alexander the Great invade Egypt” by going on a diatribe about how much war there was in the ancient world and that the ptolemaic dynasty believed themselves the rightful rulers therefore who’s to say if they did invade or just take their rightful place. There is an objective record here: whether or not people want to try and hide it behind circuitous arguments is different. If we’re going down this road I can easily redefine any known historical event with hand-wavy nonsense that doesn’t actually have anything to do with the historical record of events just “vibes.”

Re: GPT-4.5

#410
post #118

Earlier quoted context omitted.

Input price difference: 4.5 is 30x more Output price difference:4.5 is 15x more In their model evaluation scores in the appendix, 4.5 is, on average, 26% better. I don't understand the value here.

If you ran the same query set 30x or 15x on the cheaper model (and compensated for all the extra tokens the reasoning model uses), would you be able to realize the same 26% quality gain in a machine-adjudicatible kind of way?

Ignoring latency for a second, one of the tricks for boosting quality is to utilize consensus. One probability does not need to call the lesser model 30x as much to achieve these gains sorta of gains. Moreover you have to take the purported gains with a grain of salt. The models are probably trained on the evaluation sets they are benchmarked against.
Post reply on HN