I feel like OpenAI is pursuing AGI when Anthropic/Claude is pursuing making AI awesome for practical things like coding. I only ever using OpenAI's coding now as a double check against Claude. Does OpenAI have their eyes on the ball?
My usage has come down to mostly Claude (until I run out of free tier quota) and then Gemini. Claude is the best for code and Gemini 2.0 Flash is good enough while also being free (well considering how much data G has hoovered up over the years, perhaps not) and more importantly highly available. For simple queries like generating shell scripts for some plumbing, or doing some data munging, I go straight to Gemini.
GPT-4.5
271–280 of 1001 posts
Re: GPT-4.5
#272Earlier quoted context omitted.
I’m not an expert or anything, but from my vantage point, each passing release makes Altman’s confidence look more aspirational than visionary, which is a really bad place to be with that kind of money tied up. My financial manager is pretty bullish on tech so I hope he is paying close attention to the way this market space is evolving. He’s good at his job, a nice guy, and surely wears much more expensive underwear…
You're the one buying him the underwear. Don't index funds outperform managed investing? I think especially after accounting for fees, but possibly even after accounting that 50% of money managers are below average.
After the scam came to light my friend said “if I knew I was paying for those dinners, I would have been fine with Denny’s[1]”
I wanted to tell him “you would have been paying for those dinners even if he wasn’t outright stealing your money,” but that seemed insensitive so I kept my mouth shut.
0 - a local steakhouse had a portrait of this guy drawn on the wall
1 - for any non-Americans, Denny’s is a low cost diner-style restaurant.
Re: GPT-4.5
#273Earlier quoted context omitted.
I’m not an expert or anything, but from my vantage point, each passing release makes Altman’s confidence look more aspirational than visionary, which is a really bad place to be with that kind of money tied up. My financial manager is pretty bullish on tech so I hope he is paying close attention to the way this market space is evolving. He’s good at his job, a nice guy, and surely wears much more expensive underwear…
> each passing release makes Altman’s confidence look more aspirational than visionary As an LLM cynic, I feel that point passed long go, perhaps even before Altman claimed countries would start wars to conquer the territory around GPU datacenters, or promoting the dream of a 7 T-for-trillion dollar investment deal, etc. Alas, the market can remain irrational longer than I can remain solvent.
Re: GPT-4.5
#274GPT 4.5 pricing is insane: Price Input: $75.00 / 1M tokens Cached input: $37.50 / 1M tokens Output: $150.00 / 1M tokens GPT 4o pricing for comparison: Price Input: $2.50 / 1M tokens Cached input: $1.25 / 1M tokens Output: $10.00 / 1M tokens It sounds like it's so expensive and the difference in usefulness is so lacking(?) they're not even gonna keep serving it in the API for long: > GPT‑4.5 is a very large and comput…
> We look forward to learning more about its strengths, capabilities, and potential applications in real-world settings. If GPT‑4.5 delivers unique value for your use case, your feedback (opens in a new window) will play an important role in guiding our decision. "We don't really know what this is good for, but spent a lot of money and time making it and are under intense pressure to announce new things right now. If…
Re: GPT-4.5
#275My 2 cents (disclaimer: I am talking out of my ass) here is why GPTs actually suck at fluid knowledge retrievel (which is kinda their main usecase, with them being used as knowledge engines) - they've mentioned that if you train it on 'Tom Cruise was born July 3, 1962', it won't be able to answer the question "Who was born on July 3, 1962", if you don't feed it this piece of information. It can't really internally co…
So while they may have a chance of answering "What is this non mainstream novel about" they may be unable to recommend the novel since it's not a likely series of tokens in response to a request for a book recommendation.
Re: GPT-4.5
#276It is interesting that they are focusing a large part of this release on the model having a higher "EQ" (Emotional Quotient). We're far from the days of "this is not a person, we do not want to make it addictive" and getting a firm foot on the territory of "here's your new AI friend". This is very visible in the example comparing 4o with 4.5 when the user is complaining about failing a test, where 4o's response is wh…
I would like to see a humor test. So far, I have not seen any model response that has made me laugh.
Re: GPT-4.5
#277Re: GPT-4.5
#278This is such as confusing release / announcement. It seems clearly worse than Claude Sonnet 3.7, yet costs 30x as much? What's going on with their pricing? I misread it as $7.5/M input and that that was very overpriced... then realized it was 10x that much!
Re: GPT-4.5
#279GPT 4.5 pricing is insane: Price Input: $75.00 / 1M tokens Cached input: $37.50 / 1M tokens Output: $150.00 / 1M tokens GPT 4o pricing for comparison: Price Input: $2.50 / 1M tokens Cached input: $1.25 / 1M tokens Output: $10.00 / 1M tokens It sounds like it's so expensive and the difference in usefulness is so lacking(?) they're not even gonna keep serving it in the API for long: > GPT‑4.5 is a very large and comput…
Sam Altman's explanation for the restriction is a bit fluffier: https://x.com/sama/status/1895203654103351462 > bad news: it is a giant, expensive model. we really wanted to launch it to plus and pro at the same time, but we've been growing a lot and are out of GPUs. we will add tens of thousands of GPUs next week and roll it out to the plus tier then. (hundreds of thousands coming soon, and i'm pretty sure y'all wil…
Re: GPT-4.5
#280Earlier quoted context omitted.
> "Early testing shows that interacting with GPT‑4.5 feels more natural. Its broader knowledge base, improved ability to follow user intent, and greater “EQ” make it useful for tasks like improving writing, programming, and solving practical problems. We also expect it to hallucinate less." "Early testing doesn't show that it hallucinates less, but we expect that putting that sentence nearby will lead you to draw a c…
According to a graph they provide, it does hallucinate significantly less on at least one benchmark.