Live data from Hacker News

GPT-4.5

openai.com

271–280 of 1001 posts

Re: GPT-4.5

#271

I feel like OpenAI is pursuing AGI when Anthropic/Claude is pursuing making AI awesome for practical things like coding. I only ever using OpenAI's coding now as a double check against Claude. Does OpenAI have their eyes on the ball?

My usage has come down to mostly Claude (until I run out of free tier quota) and then Gemini. Claude is the best for code and Gemini 2.0 Flash is good enough while also being free (well considering how much data G has hoovered up over the years, perhaps not) and more importantly highly available. For simple queries like generating shell scripts for some plumbing, or doing some data munging, I go straight to Gemini.

Claude still can't make real time web searches though for RAG, unless I'm missing something.

Re: GPT-4.5

#272
post #177

Earlier quoted context omitted.

I’m not an expert or anything, but from my vantage point, each passing release makes Altman’s confidence look more aspirational than visionary, which is a really bad place to be with that kind of money tied up. My financial manager is pretty bullish on tech so I hope he is paying close attention to the way this market space is evolving. He’s good at his job, a nice guy, and surely wears much more expensive underwear…

You're the one buying him the underwear. Don't index funds outperform managed investing? I think especially after accounting for fees, but possibly even after accounting that 50% of money managers are below average.

A friend got taken in by a Ponzi scheme operator several years ago. The guy running it was known for taking his clients out to lavish dinners and events all the time.[0]

After the scam came to light my friend said “if I knew I was paying for those dinners, I would have been fine with Denny’s[1]”

I wanted to tell him “you would have been paying for those dinners even if he wasn’t outright stealing your money,” but that seemed insensitive so I kept my mouth shut.

0 - a local steakhouse had a portrait of this guy drawn on the wall

1 - for any non-Americans, Denny’s is a low cost diner-style restaurant.

Re: GPT-4.5

#273
post #234

Earlier quoted context omitted.

I’m not an expert or anything, but from my vantage point, each passing release makes Altman’s confidence look more aspirational than visionary, which is a really bad place to be with that kind of money tied up. My financial manager is pretty bullish on tech so I hope he is paying close attention to the way this market space is evolving. He’s good at his job, a nice guy, and surely wears much more expensive underwear…

> each passing release makes Altman’s confidence look more aspirational than visionary As an LLM cynic, I feel that point passed long go, perhaps even before Altman claimed countries would start wars to conquer the territory around GPU datacenters, or promoting the dream of a 7 T-for-trillion dollar investment deal, etc. Alas, the market can remain irrational longer than I can remain solvent.

That $7 trillion dollar ask pushed me from skeptical to full-on eye-roll emoji land— the dude is clearly a narcissist with delusions of grandeur— but it’s getting worse. Considering the $200 pro subscription was significantly unprofitable before this model came out, imagine how astonishingly expensive this model must be to run at many times that price.

Re: GPT-4.5

#274
post #26

GPT 4.5 pricing is insane: Price Input: $75.00 / 1M tokens Cached input: $37.50 / 1M tokens Output: $150.00 / 1M tokens GPT 4o pricing for comparison: Price Input: $2.50 / 1M tokens Cached input: $1.25 / 1M tokens Output: $10.00 / 1M tokens It sounds like it's so expensive and the difference in usefulness is so lacking(?) they're not even gonna keep serving it in the API for long: > GPT‑4.5 is a very large and comput…

> We look forward to learning more about its strengths, capabilities, and potential applications in real-world settings. If GPT‑4.5 delivers unique value for your use case, your feedback (opens in a new window) will play an important role in guiding our decision. "We don't really know what this is good for, but spent a lot of money and time making it and are under intense pressure to announce new things right now. If…

it's so over, pretraining is ngmi. maybe sam Altman was wrong after all ? https://www.lycee.ai/blog/why-sam-altman-is-wrong

Re: GPT-4.5

#275

My 2 cents (disclaimer: I am talking out of my ass) here is why GPTs actually suck at fluid knowledge retrievel (which is kinda their main usecase, with them being used as knowledge engines) - they've mentioned that if you train it on 'Tom Cruise was born July 3, 1962', it won't be able to answer the question "Who was born on July 3, 1962", if you don't feed it this piece of information. It can't really internally co…

You make a good point: I think these LLM's have a strong bias towards recommending the most popular things in pop culture since they really only find the most likely tokens and report on that.

So while they may have a chance of answering "What is this non mainstream novel about" they may be unable to recommend the novel since it's not a likely series of tokens in response to a request for a book recommendation.

Re: GPT-4.5

#276
post #73

It is interesting that they are focusing a large part of this release on the model having a higher "EQ" (Emotional Quotient). We're far from the days of "this is not a person, we do not want to make it addictive" and getting a firm foot on the territory of "here's your new AI friend". This is very visible in the example comparing 4o with 4.5 when the user is complaining about failing a test, where 4o's response is wh…

I would like to see a humor test. So far, I have not seen any model response that has made me laugh.

How does the following stand-up routine by Claude 3.7 Sonnet work for you?

https://gally.net/temp/20250225claudestandup2.html

Re: GPT-4.5

#278

This is such as confusing release / announcement. It seems clearly worse than Claude Sonnet 3.7, yet costs 30x as much? What's going on with their pricing? I misread it as $7.5/M input and that that was very overpriced... then realized it was 10x that much!

Is it worse than clause sonnet with reasoning enabled or disabled ?

Re: GPT-4.5

#279
post #26

GPT 4.5 pricing is insane: Price Input: $75.00 / 1M tokens Cached input: $37.50 / 1M tokens Output: $150.00 / 1M tokens GPT 4o pricing for comparison: Price Input: $2.50 / 1M tokens Cached input: $1.25 / 1M tokens Output: $10.00 / 1M tokens It sounds like it's so expensive and the difference in usefulness is so lacking(?) they're not even gonna keep serving it in the API for long: > GPT‑4.5 is a very large and comput…

Sam Altman's explanation for the restriction is a bit fluffier: https://x.com/sama/status/1895203654103351462 > bad news: it is a giant, expensive model. we really wanted to launch it to plus and pro at the same time, but we've been growing a lot and are out of GPUs. we will add tens of thousands of GPUs next week and roll it out to the plus tier then. (hundreds of thousands coming soon, and i'm pretty sure y'all wil…

Bad news: Sam Altman runs the show.

Re: GPT-4.5

#280

Earlier quoted context omitted.

> "Early testing shows that interacting with GPT‑4.5 feels more natural. Its broader knowledge base, improved ability to follow user intent, and greater “EQ” make it useful for tasks like improving writing, programming, and solving practical problems. We also expect it to hallucinate less." "Early testing doesn't show that it hallucinates less, but we expect that putting that sentence nearby will lead you to draw a c…

According to a graph they provide, it does hallucinate significantly less on at least one benchmark.

It hallucinates at 37% on SimpleQA yeah, which is a set of very difficult questions inviting hallucinations. Claude 3.5 Sonnet (the June 2024 editiom, before October update and before 3.7) hallucinated at 35%. I think this is more of an indication of how behind OpenAI has been in this area.
Post reply on HN