Live data from Hacker News

GPT-4.5

openai.com

771–780 of 1001 posts

Re: GPT-4.5

#771

Earlier quoted context omitted.

I think it's fairer to compare it to the original GPT-4 which might the equivalent in term of "size" (though we don't have actual numbers for either). GPT-4: Input $30.00 / 1M tokens ; Output $60.00 / 1M tokens So 4.5 is 2.5x more expensive. I think they announced this as their last non-reasoning model, so it was maybe with the goal of stretching pre-training as far as they could, just to see what new capabilities wo…

Why would that be fairer? We can assume they did incorporate all learnings and optimizations they made post gpt-4 launch, no?

Definitely not. They don't distill their original models. 4o is a much more distilled and cheaper version of 4. I assume 4.5o would be a distilled and cheaper version of 4.5.

It'd be weird to release a distilled version without ever releasing the base undistilled version.

Re: GPT-4.5

#772

Earlier quoted context omitted.

AI as it stands in 2025 is an amazing technology, but it is not a product at all . As a result, OpenAI simply does not have a business model, even if they are trying to convince the world that they do. My bet is that they're currently burning through other people's capital at an amazing rate, but that they are light-years from profitability They are also being chased by fierce competition and OpenSource which is very…

> AI as it stands in 2025 is an amazing technology, but it is not a product at all. Here I'm assuming "AI" to mean what's broadly called Generative AI (LLMs, photo, video generation) I genuinely am struggling to see what the product is too. The code assistant use cases are really impressive across the board (and I'm someone who was vocally against them less than a year ago), and I pay for Github CoPilot (for now) but…

> I genuinely am struggling to see what the product is too.

They're nice for summarizing and categorizing text. We've had good solutions for that before, too (BERT, et al), but LLM's are marginally nicer.

> Is there a market of people clamoring to use/get anything GenAI related?

No. LLM's are lame and uncool. Kids especially dislike them a lot on that basis alone.

Re: GPT-4.5

#773
I want less and less of these "do it all models", what I want is specific models for the exact task I need.

Then what I want is a platform with a generic AI on top that can pick the correct expert models based on what I asked it to do.

Kinda what Apple is attempting with their Small Language Model thing?

Re: GPT-4.5

#774
post #433

Earlier quoted context omitted.

My guess is that you're right about that being what's next (or maybe almost next) from them, but I think they'll save the name GPT-5 for the next actually-trained model (like 4.5 but a bigger jump), and use a different kind of name for the routing model. Even by their poor standards at naming it would be weird to introduce a completely new type/concept, that can loop in models including the 4 / 4.5 series, while nami…

They already confirmed GPT-5 will be a unified model "months" away. Elsewhere they claimed that it will not just be a router but a "unified" model. https://www.theverge.com/news/611365/openai-gpt-4-5-roadmap-...

Interesting, thanks for sharing - definitely makes me withdraw my confidence in that prediction, though I still think there's a decent chance they change their mind about that as it seems to me like an even worse naming decision than their previous shit name choices!

Re: GPT-4.5

#775
post #280

Earlier quoted context omitted.

According to a graph they provide, it does hallucinate significantly less on at least one benchmark.

It hallucinates at 37% on SimpleQA yeah, which is a set of very difficult questions inviting hallucinations. Claude 3.5 Sonnet (the June 2024 editiom, before October update and before 3.7) hallucinated at 35%. I think this is more of an indication of how behind OpenAI has been in this area.

I wonder how it's even possible to evaluate this kind of thing without data leakage. Correct answers to specific, factual questions are only possible if the model has seen those answers in the training data, so how reliable can the benchmark be if the test dataset is contaminated with training data?

Or is the assumption that the training set is so big it doesn't matter?

Re: GPT-4.5

#777
post #26

GPT 4.5 pricing is insane: Price Input: $75.00 / 1M tokens Cached input: $37.50 / 1M tokens Output: $150.00 / 1M tokens GPT 4o pricing for comparison: Price Input: $2.50 / 1M tokens Cached input: $1.25 / 1M tokens Output: $10.00 / 1M tokens It sounds like it's so expensive and the difference in usefulness is so lacking(?) they're not even gonna keep serving it in the API for long: > GPT‑4.5 is a very large and comput…

usefulness is bound to scope/purpose, even if innovation stops, in 3y (thanks to hw and tuning progress ) when 4o costs 0.1$/M and 4.5 1$/M even being a small improvement ( which is not imo ), you will chose to use 4.5 , exactly like no one now want to use 3.5

Re: GPT-4.5

#778
post #26

GPT 4.5 pricing is insane: Price Input: $75.00 / 1M tokens Cached input: $37.50 / 1M tokens Output: $150.00 / 1M tokens GPT 4o pricing for comparison: Price Input: $2.50 / 1M tokens Cached input: $1.25 / 1M tokens Output: $10.00 / 1M tokens It sounds like it's so expensive and the difference in usefulness is so lacking(?) they're not even gonna keep serving it in the API for long: > GPT‑4.5 is a very large and comput…

It's priced like this because it can generate erotica.

Re: GPT-4.5

#779

Oh this makes sense. chatGPT results have taken a nose dive in quality lately. It couldn't write a simple rename function for me yesterday, still buggy after seven attempts. I'm more and more convinced that they dumb down the core product when they plan to release a new version to make the difference seem bigger.

>It couldn't write a simple rename function for me yesterday, still buggy after seven attempts. I'm surprised and a bit nervous about that. We intend to bootstrap a large project with it!! Both ChatGPT 4o (fast) and ChatGPT o1 (a bit slower, deeper thinking) should easily be able to do this without fail. Where did it go wrong? Could you please link to your chat? About my project: I run the sovereign State of Utopia (…

with 4.5? Why? It's only meant for creative writing.

Re: GPT-4.5

#780

My 2 cents (disclaimer: I am talking out of my ass) here is why GPTs actually suck at fluid knowledge retrievel (which is kinda their main usecase, with them being used as knowledge engines) - they've mentioned that if you train it on 'Tom Cruise was born July 3, 1962', it won't be able to answer the question "Who was born on July 3, 1962", if you don't feed it this piece of information. It can't really internally co…

You make a good point: I think these LLM's have a strong bias towards recommending the most popular things in pop culture since they really only find the most likely tokens and report on that. So while they may have a chance of answering "What is this non mainstream novel about" they may be unable to recommend the novel since it's not a likely series of tokens in response to a request for a book recommendation.

That's really interesting - just made me think about some AI guy at Twitter (when it was called that) talking about how hard it is to create a recommender system that doesn't just flood everyone with what's popular righr now. Since LLMs are neural networks as well, maybe the recommendation algorithms they learn suffer from the same issues
Post reply on HN