Live data from Hacker News

GPT-4.5

openai.com

451–460 of 1001 posts

Re: GPT-4.5

#451
post #26

GPT 4.5 pricing is insane: Price Input: $75.00 / 1M tokens Cached input: $37.50 / 1M tokens Output: $150.00 / 1M tokens GPT 4o pricing for comparison: Price Input: $2.50 / 1M tokens Cached input: $1.25 / 1M tokens Output: $10.00 / 1M tokens It sounds like it's so expensive and the difference in usefulness is so lacking(?) they're not even gonna keep serving it in the API for long: > GPT‑4.5 is a very large and comput…

The price really is eye watering. At a glance, my first impression is this is something like Llama 3.1 405B, where the primary value may be realized in generating high quality synthetic data for training rather than direct use. I keep a little google spreadsheet with some charts to help visualize the landscape at a glance in terms of capability/price/throughput, bringing in the various index scores as they become ava…

> feel free to copy and claim as your own.

That's a nice sentiment, but I'd encourage you to add a license or something. The basic "something" would be adding a canonical URL into the spreadsheet itself somewhere, along with a notification that users can do what they want other than removing that URL. (And the URL would be described as "the original source" or something, not a claim that the particular version/incarnation someone is looking at is the same as what is at that URL.)

The risk is that someone will accidentally introduce errors or unsupportable claims, and people with the modified spreadsheet won't know that it's not The spreadsheet and so will discount its accuracy or trustability. (If people are trying to deceive others into thinking it's the original, they'll remove the notice, but that's a different problem.) It would be a shame for people to lose faith in your work because of crap that other people do that you have no say in.

Re: GPT-4.5

#452
post #342

I got gpt-4.5-preview to summarize this discussion thread so far (at 324 comments): hn-summary.sh 43197872 -m gpt-4.5-preview Using this script: https://til.simonwillison.net/llms/claude-hacker-news-themes... Here's the result: https://gist.github.com/simonw/5e9f5e94ac8840f698c280293d399... It took 25797 input tokens and 1225 input tokens, for a total cost (calculated using https://tools.simonwillison.net/llm-prices…

Didn't seem to realize that "Still more coherent than the OpenAI lineup" wouldn't make sense out of context. (The actual comment quoted there is responding to someone who says they'd name their models Foo, Bar, Baz.)

Re: GPT-4.5

#453
post #24

Considering both this blog post and the livestream demos, I am underwhelmed. Having just finished the stream, I had a real "was that all" moment, which on one hand shows how spoiled I've gotten by new models impressing me, but on another feels like OpenAI really struggles to stay ahead of their competitors. What has been shown feels like it could be achieved using a custom system prompt on older versions of OpenAIs m…

Funny you should suggest that it seems like a revised system prompt: https://chatgpt.com/share/67c0fda8-a940-800f-bbdc-6674a8375f...

Re: GPT-4.5

#454

Earlier quoted context omitted.

What it confirms, I think, is, that we are going to need a lot more chips.

Further confirmation, IMO, that the idea that any of this leads to anything close to AGI is people getting high on their own supply (in some cases literally). LLMs are a great tool for what is effectively collected knowledge search and summary (so long as you are willing to accept that you have to verify all of the 'knowledge' they spit back because they always have the ability to go off the rails) but they have been…

Well said. 100% agree

Re: GPT-4.5

#455

It is interesting that they are focusing a large part of this release on the model having a higher "EQ" (Emotional Quotient). We're far from the days of "this is not a person, we do not want to make it addictive" and getting a firm foot on the territory of "here's your new AI friend". This is very visible in the example comparing 4o with 4.5 when the user is complaining about failing a test, where 4o's response is wh…

Anthropic pretty much abandoned this direction after Claude 3, and said it wasn't what they wanted [1]. Claude 3.5+ is extremely dry and neutral, it doesn't seem to have the same training. >Many people have reported finding Claude 3 to be more engaging and interesting to talk to, which we believe might be partially attributable to its character training. This wasn’t the core goal of character training, however. Model…

It's the opposite incentive to ad-funded social media. One wants to drain your wallet and keep you hooked, the other wants you to spend as little of their funding as possible finding what you're looking for.

Re: GPT-4.5

#456

Seeing OpenAI and Anthropic go different routes here is interesting. It is worth moving past the initial knee jerk reaction of this model being unimpressive and some of the comments about "they spent a massive amount of money and had to ship something for it..." * Anthropic appears to be making a bet that a single paradigm (reasoning) can create a model which is excellent for all use cases. * OpenAI seems to be betti…

It can never be just reasoning, right? Reasoning is the multiplier on some base model, and surely no amount of reasoning on top of something like gpt-2 will get you o1. This model is too expensive right now, but as compute gets cheaper — and we have to keep in mind, that it will — having a better base to multiply with will enable things that just more thinking won't.

You can try for yourself with the distilled R1's that Deepseek released. The qwen-7b based model is quite impressive for its size and it can do a lot with additional context provided. I imagine for some domains you can provide enough context and let the inference time eventually solve it, for others you can't.

Re: GPT-4.5

#457

Earlier quoted context omitted.

Maybe if they build a few more data centers, they'll be able to construct their machine god. Just a few more dedicated power plants, a lake or two, a few hundred billion more and they'll crack this thing wide open. And maybe Tesla is going to deliver truly full self driving tech any day now. And Star Citizen will prove to have been worth it along along, and Bitcoin will rain from the heavens. It's very difficult to r…

Star Citizen is a working model of how to do UBI. That entire staff of a thousand people is the test case.

Finally, someone gets it.

Re: GPT-4.5

#458

Honestly, the most astounding part of this announcement is their comparison to o3-mini with QA prompts. EIGHTY PERCENT hallucination rate? Are you kidding me? I get that the model is meant to be used for logic and reasoning, but nowhere does OpenAI make this explicitly clear. A majority of users are going to be thinking, "oh newer is better," and pick that.

Very nice catch, I was under the impression that o3-mini was "as good" as o1 on all dimensions. Seems the takeaway is that any form of quantization/distillation ends up hurting factual accuracy (but not reasoning performance), and there are diminishing returns to reducing hallucinations by model-scaling or RLHF'ing. I guess then that other approaches are needed to achieve single-digit "hallucination" rates. All of wikipedia compresses down to < 50GB though, so it's not immediately clear that you can't have good factual accuracy with a small sparse model

Re: GPT-4.5

#459

Earlier quoted context omitted.

> We don't really know what this is good for Oh come on. Think how long of a gap there was between the first microcomputer and VisiCalc. Or between the start of the internet and social networking. First of all, it's going to take us 10 years to figure out how to use LLM's to their full productive potential. And second of all, it's going to take us collectively a long time to also figure out how much accuracy is neces…

The Internet had plenty of very productive use cases before social networking, even from its most nascent origins. Spending billions building something on the assumption that someone else will figure out what it's good for, is not good business.

It's incredibly good and lucrative business. You are confusing scientifically sound and well-planned out and conservative risk tolerance with good business

Re: GPT-4.5

#460
post #372

Earlier quoted context omitted.

From Sam's twitter: > After that, a top goal for us is to unify o-series models and GPT-series models by creating systems that can use all our tools, know when to think for a long time or not, and generally be useful for a very wide range of tasks. > In both ChatGPT and our API, we will release GPT-5 as a system that integrates a lot of our technology, including o3. We will no longer ship o3 as a standalone model. Yo…

I worry eliminating consumer choice will drive up prices for only a nominal gain in utility for most users.

I'm more worries they'll push down their costs by making it harder to get the reasoning models to run, but either would suck.
Post reply on HN