Live data from Hacker News

OpenAI dropped the price of o3 by 80%

twitter.com

31–40 of 518 posts

Re: OpenAI dropped the price of o3 by 80%

#31

how do we know it's not a quantized version of o3? what's stopping these firms from announcing the full model to perform well on the benchmarks and then gradually quantizing it (first at Q8 so no one notices, then Q6, then Q4, ...). I have a suspicion that's how they were able to get gpt-4-turbo so fast. In practice, I found it inferior to the original GPT-4 but the company probably benchmaxxed the hell out of the tu…

Quantization is a massive efficiency gain for near negligible drop in quality. If the tradeoff is quantization for an 80 percent price drop I would take that any day of the week.

You may be right that the tradeoff is worth it, but it should be advertised as such. You shouldn't think you're paying for full o3, even if they're heavily discounting it.

Re: OpenAI dropped the price of o3 by 80%

#33

how do we know it's not a quantized version of o3? what's stopping these firms from announcing the full model to perform well on the benchmarks and then gradually quantizing it (first at Q8 so no one notices, then Q6, then Q4, ...). I have a suspicion that's how they were able to get gpt-4-turbo so fast. In practice, I found it inferior to the original GPT-4 but the company probably benchmaxxed the hell out of the tu…

I swear every time a new model is released it's great at first but then performance gets worse over time. I figured they were fine-tuning it to get rid of bad output which also nerfed the really good output. Now I'm wondering if they were quantizing it.

It seems that least Google is overselling their compute capacity.

You pay monthly fee, but Gemini is completely jammed 5-6 hours when North America is working.

Re: OpenAI dropped the price of o3 by 80%

#34
post #29

always seemed to me that efficient caching strategies could greatly reduce costs… wonder if they cooked up something new

How are LLMs cached? Every prompt would be different so it's not clear how that would work. Unless you're talking about caching the model weights...

A lot of the prompt is always the same: the instructions, the context, the codebase (if you are coding), etc.

Re: OpenAI dropped the price of o3 by 80%

#35
post #28

You know. because LLMs can only be built by corporations... but because they're so easy to build, I see the price going down massively thanks to competition. Consumers benefit because all the companies are trying to out run each other.

And then they all go out of business, since models cost a fortune to build, and their fan club is left staring at their computers trying to remember how to do anything without getting it served on a silver plate.

Investors pouring money, its probably impossible to go out of business, at least for the big ones, until investors realise this is wrong hill to die on.

Re: OpenAI dropped the price of o3 by 80%

#37
post #27

It's going to be a race to the bottom, they have no moat.

Especially now that they are second in the race (behind Anthropic) and lot of free-to-download and free-to-use models are now starting to be viable competitors. Once new MacBooks and iPhones have enough memory onboard this is going to be a disaster for OpenAI and other providers.

I'm not sure they're scared of Anthropic - they're doing great work but afaict running into some scaling issues and really focused on winning over developers at the moment.

If I was OpenAI (or Anthropic for that matter) I would remain scared of Google, who is now awake and able to dump Gemini 2.5 pro on the market at costs that I'm not sure people without their own hardware can compete with, and with the infrastructure to handle everyone switching to them tomorrow.

Re: OpenAI dropped the price of o3 by 80%

#38
post #29

always seemed to me that efficient caching strategies could greatly reduce costs… wonder if they cooked up something new

How are LLMs cached? Every prompt would be different so it's not clear how that would work. Unless you're talking about caching the model weights...

You would use a KV cache to cache a significant chunk of the inference work.

Re: OpenAI dropped the price of o3 by 80%

#39
post #27

It's going to be a race to the bottom, they have no moat.

Especially now that they are second in the race (behind Anthropic) and lot of free-to-download and free-to-use models are now starting to be viable competitors. Once new MacBooks and iPhones have enough memory onboard this is going to be a disaster for OpenAI and other providers.

What do you mean, Google is number 1

Re: OpenAI dropped the price of o3 by 80%

#40
I don't know if this is OpenAI's intention, but the little message "you've reached your usage limit!" is actively disincentivizing me from subscribing. For my purposes, the free model is more than good enough; the difference before and after is negligible. I honestly wouldn't pay a dollar.

That said, I'm absolutely willing to hear people out on "value-adds" I am missing out on; I'm not a knee-jerk hater (For context, I work with large, complex & private databases/platforms, so its not really possible for me to do anything but ask for scripting suggestions).

Also, I am 100% expecting a sad day when I'll be forced to subscribe, unless I want to read dick pill ads shoehorned in to the answers (looking at you, YouTube). I do worry about getting dependent on this tool and watching it become enshittified.

Post reply on HN