Live data from Hacker News

GPT-4.5

openai.com

221–230 of 1001 posts

Re: GPT-4.5

#221

Earlier quoted context omitted.

> We look forward to learning more about its strengths, capabilities, and potential applications in real-world settings. If GPT‑4.5 delivers unique value for your use case, your feedback (opens in a new window) will play an important role in guiding our decision. "We don't really know what this is good for, but spent a lot of money and time making it and are under intense pressure to announce new things right now. If…

> "Early testing shows that interacting with GPT‑4.5 feels more natural. Its broader knowledge base, improved ability to follow user intent, and greater “EQ” make it useful for tasks like improving writing, programming, and solving practical problems. We also expect it to hallucinate less." "Early testing doesn't show that it hallucinates less, but we expect that putting that sentence nearby will lead you to draw a c…

According to a graph they provide, it does hallucinate significantly less on at least one benchmark.

Re: GPT-4.5

#222

Earlier quoted context omitted.

Altman's claim and NVIDIA's consumer launch supply problems may be related - OpenAI may be eating up the GPU supply...

OpenAI is not purchasing consumer 5090s... :)

No, but the supply constraints are part of what is driving the insane prices. Every chip they use for consumer grade instead of commercial grade is a potential loss of potential income.

Re: GPT-4.5

#224

I feel like OpenAI is pursuing AGI when Anthropic/Claude is pursuing making AI awesome for practical things like coding. I only ever using OpenAI's coding now as a double check against Claude. Does OpenAI have their eyes on the ball?

My usage has come down to mostly Claude (until I run out of free tier quota) and then Gemini. Claude is the best for code and Gemini 2.0 Flash is good enough while also being free (well considering how much data G has hoovered up over the years, perhaps not) and more importantly highly available. For simple queries like generating shell scripts for some plumbing, or doing some data munging, I go straight to Gemini.

> My usage has come down to mostly Claude (until I run out of free tier quota) and then Gemini

Yep, exactly same here.

Gemini 2.0 Flash is extremely good, and I've yet to hit any usage limits with them - for heavy usage I just go to Gemini directly. For "talk to an expert" usage, Claude is hard to beat though.

Re: GPT-4.5

#225
post #26

GPT 4.5 pricing is insane: Price Input: $75.00 / 1M tokens Cached input: $37.50 / 1M tokens Output: $150.00 / 1M tokens GPT 4o pricing for comparison: Price Input: $2.50 / 1M tokens Cached input: $1.25 / 1M tokens Output: $10.00 / 1M tokens It sounds like it's so expensive and the difference in usefulness is so lacking(?) they're not even gonna keep serving it in the API for long: > GPT‑4.5 is a very large and comput…

Input price difference: 4.5 is 30x more Output price difference:4.5 is 15x more In their model evaluation scores in the appendix, 4.5 is, on average, 26% better. I don't understand the value here.

Einstein's IQ = 3.5x chimpanzees IQs, right?

Re: GPT-4.5

#226

I'm one week in on heavy grok usage. I didn't think I'd say this, but for personal use, I'm considering cancelling my OpenAI plan. The one thing I wish grok had was more separation of the UI from X itself. The interface being so coupled to X puts me off and makes it feel like a second-hand citizen. I like ChatGPTs minimalist UI.

Theres grok.com which is standalone and with its own UI

There's also a standalone Grok app at least on iOS.

Re: GPT-4.5

#227
brief and detailed summaries by chatgpt (4o):

Brief Summary (40-50 words)

OpenAI’s GPT-4.5 is a research preview of their most advanced language model yet, emphasizing improved pattern recognition, creativity, and reduced hallucinations. It enhances unsupervised learning, has better emotional intelligence, and excels in writing, programming, and problem-solving. Available for ChatGPT Pro users, it also integrates into APIs for developers.

Detailed Summary (200 words)

OpenAI has introduced *GPT-4.5*, a research preview of its most advanced language model, focusing on *scaling unsupervised learning* to enhance pattern recognition, knowledge depth, and reliability. It surpasses previous models in *natural conversation, emotional intelligence (EQ), and nuanced understanding of user intent*, making it particularly useful for writing, programming, and creative tasks.

GPT-4.5 benefits from *scalable training techniques* that improve its steerability and ability to comprehend complex prompts. Compared to GPT-4o, it has a *higher factual accuracy and lower hallucination rates*, making it more dependable across various domains. While it does not employ reasoning-based pre-processing like OpenAI o1, it complements such models by excelling in general intelligence.

Safety improvements include *new supervision techniques* alongside traditional reinforcement learning from human feedback (RLHF). OpenAI has tested GPT-4.5 under its *Preparedness Framework* to ensure alignment and risk mitigation.

*Availability*: GPT-4.5 is accessible to *ChatGPT Pro users*, rolling out to other tiers soon. Developers can also use it in *Chat Completions API, Assistants API, and Batch API*, with *function calling and vision capabilities*. However, it remains computationally expensive, and OpenAI is evaluating its long-term API availability.

GPT-4.5 represents a *major step in AI model scaling*, offering *greater creativity, contextual awareness, and collaboration potential*.

Re: GPT-4.5

#228
post #73

It is interesting that they are focusing a large part of this release on the model having a higher "EQ" (Emotional Quotient). We're far from the days of "this is not a person, we do not want to make it addictive" and getting a firm foot on the territory of "here's your new AI friend". This is very visible in the example comparing 4o with 4.5 when the user is complaining about failing a test, where 4o's response is wh…

I would like to see a humor test. So far, I have not seen any model response that has made me laugh.

My benchmark for this has been asking the model to write some tweets in the style of dril, a popular user who writes short funny tweets. Sometimes I include a few example tweets in the prompt too. Here's an example of results I got from Claude 3 Opus and GPT 4 for this last year: https://bsky.app/profile/macil.tech/post/3kpcvicmirs2v. My opinion is that Claude's results were mostly bangers while GPT's were all a bit groanworthy. I need to try this again with the latest models sometime.

Re: GPT-4.5

#229
post #26

GPT 4.5 pricing is insane: Price Input: $75.00 / 1M tokens Cached input: $37.50 / 1M tokens Output: $150.00 / 1M tokens GPT 4o pricing for comparison: Price Input: $2.50 / 1M tokens Cached input: $1.25 / 1M tokens Output: $10.00 / 1M tokens It sounds like it's so expensive and the difference in usefulness is so lacking(?) they're not even gonna keep serving it in the API for long: > GPT‑4.5 is a very large and comput…

The price really is eye watering. At a glance, my first impression is this is something like Llama 3.1 405B, where the primary value may be realized in generating high quality synthetic data for training rather than direct use. I keep a little google spreadsheet with some charts to help visualize the landscape at a glance in terms of capability/price/throughput, bringing in the various index scores as they become ava…

very impressive... also interested in your trip planner, it looks like invite only at the moment, but... would it be rude to ask for an invite?

Re: GPT-4.5

#230
post #26

GPT 4.5 pricing is insane: Price Input: $75.00 / 1M tokens Cached input: $37.50 / 1M tokens Output: $150.00 / 1M tokens GPT 4o pricing for comparison: Price Input: $2.50 / 1M tokens Cached input: $1.25 / 1M tokens Output: $10.00 / 1M tokens It sounds like it's so expensive and the difference in usefulness is so lacking(?) they're not even gonna keep serving it in the API for long: > GPT‑4.5 is a very large and comput…

Now the real question about AI automation starts. Is it cheaper to pay a human to do the task or a AI company?

It still not smart enough to replace for example customer service.
Post reply on HN