Earlier quoted context omitted.
I mean sure, it's very promising if OpenAI's future is your only metric. It gets notably darker if you look at the broader picture of ChatGPT (and company)'s impact on our society. * We have people uploading tons of zero-effort slop pieces to all manner of online storefronts, and making people less likely to buy overall because they assume everything is AI now * We have an uncomfortable community of, to be blunt, act…
Dying for a reference on the cult stuff, a quick search didn’t provide anything interesting.
OpenAI dropped the price of o3 by 80%
111–120 of 518 posts
Re: OpenAI dropped the price of o3 by 80%
#112Earlier quoted context omitted.
I've heard lots of people say that, but no objective reproducible benchmarks confirm such a thing happening often. Could this simply be a case of novelty/excitement for a new model fading away as you learn more about its shortcomings?
there's definitely measurements (eg https://hdsr.mitpress.mit.edu/pub/y95zitmz/release/2 ) but I imagine they're rare because those benchmarks are expensive, so nobody keeps running them all the time? Anecdotally, it's quite clear that some models are throttled during the day (eg Claude sometimes falls back to "concise mode" - with and without a warning on the app). You can tell if you're using Windsurf/Cursor too -…
Re: OpenAI dropped the price of o3 by 80%
#113Re: OpenAI dropped the price of o3 by 80%
#114how do we know it's not a quantized version of o3? what's stopping these firms from announcing the full model to perform well on the benchmarks and then gradually quantizing it (first at Q8 so no one notices, then Q6, then Q4, ...). I have a suspicion that's how they were able to get gpt-4-turbo so fast. In practice, I found it inferior to the original GPT-4 but the company probably benchmaxxed the hell out of the tu…
Re: OpenAI dropped the price of o3 by 80%
#115Earlier quoted context omitted.
It seems that least Google is overselling their compute capacity. You pay monthly fee, but Gemini is completely jammed 5-6 hours when North America is working.
Gemini is simply that good. I’m trying out Claude 4 every now and then and go back to Gemini to fix its mess…
Re: OpenAI dropped the price of o3 by 80%
#116Earlier quoted context omitted.
there's definitely measurements (eg https://hdsr.mitpress.mit.edu/pub/y95zitmz/release/2 ) but I imagine they're rare because those benchmarks are expensive, so nobody keeps running them all the time? Anecdotally, it's quite clear that some models are throttled during the day (eg Claude sometimes falls back to "concise mode" - with and without a warning on the app). You can tell if you're using Windsurf/Cursor too -…
I feel this too. I swear some of the coding Claude Code does on weekends is superior to the weekdays. It just has these eureka moments every now and then.
Trusting these LLM providers today is as risky as trusting Facebook as a platform, when they were pushing their “opensocial” stuff
Re: OpenAI dropped the price of o3 by 80%
#117Despite the popular take that LLMs have no moat and are burning cash, I find OpenAI's situation really promising. Just yesterday, they reported an annualized revenue run rate of 10B. Their last funding round in March valued them at 300B. Despite losing 5B last year, they are growing really fast - 30x revenue with over 500M active users. It reminds me a lot of Uber in its earlier years—fast growth, heavy investment, b…
The problem is your costs also scale with revenue. Ideally you want to have control costs as you scale (the first you build is expensive, but as you make more your costs come down). For OpenAI, the more people use the product, the same you spend on compute unless they can supplement it with another ways of generating revenue. I dont unfortunately think OpenAI will be able to hit sustained profitability (see Netflix f…
Obviously, lots of nerds on HN have preferences for Gemini and Claude, and having used all three I completely get why that is. But we should remember we're not representative of the whole addressable market. There were probably nerds on like ancient dial-up bulletin boards explaining why Betamax was going to win, too.
Re: OpenAI dropped the price of o3 by 80%
#118Re: OpenAI dropped the price of o3 by 80%
#119how do we know it's not a quantized version of o3? what's stopping these firms from announcing the full model to perform well on the benchmarks and then gradually quantizing it (first at Q8 so no one notices, then Q6, then Q4, ...). I have a suspicion that's how they were able to get gpt-4-turbo so fast. In practice, I found it inferior to the original GPT-4 but the company probably benchmaxxed the hell out of the tu…
Re: OpenAI dropped the price of o3 by 80%
#120Earlier quoted context omitted.
I've heard lots of people say that, but no objective reproducible benchmarks confirm such a thing happening often. Could this simply be a case of novelty/excitement for a new model fading away as you learn more about its shortcomings?
there's definitely measurements (eg https://hdsr.mitpress.mit.edu/pub/y95zitmz/release/2 ) but I imagine they're rare because those benchmarks are expensive, so nobody keeps running them all the time? Anecdotally, it's quite clear that some models are throttled during the day (eg Claude sometimes falls back to "concise mode" - with and without a warning on the app). You can tell if you're using Windsurf/Cursor too -…
You've also made the mistake of conflating what's served via API platforms which are meant to be stable, and frontends which have no stability guarantees, and are very much iterated on in terms of the underlying model and system prompts. The GPT-4o sycophancy debacle was only on the specific model that's served via the ChatGPT frontend and never impacted the stable snapshots on the API.
I have never seen any sort of compelling evidence that any of the large labs tinkers with their stable, versioned model releases that are served via their API platforms.