Live data from Hacker News

OpenAI dropped the price of o3 by 80%

twitter.com

21–30 of 518 posts

Re: OpenAI dropped the price of o3 by 80%

#21

Note that they have not actually dropped the price yet: https://x.com/OpenAIDevs/status/1932463601119637532 > We’ll post to @openaidevs once the new pricing is in full effect. In $10… 9… 8… There is also speculation that they are only dropping the input price, not the output price (which includes the reasoning tokens).

I think that was a joke. New pricing is already in place: Input: $2.00 / 1M tokens Cached input: $0.50 / 1M tokens Output: $8.00 / 1M tokens https://openai.com/api/pricing/ Now cheaper than gpt-4o and same price as gpt-4.1 (!).

It is slower though

Re: OpenAI dropped the price of o3 by 80%

#22

how do we know it's not a quantized version of o3? what's stopping these firms from announcing the full model to perform well on the benchmarks and then gradually quantizing it (first at Q8 so no one notices, then Q6, then Q4, ...). I have a suspicion that's how they were able to get gpt-4-turbo so fast. In practice, I found it inferior to the original GPT-4 but the company probably benchmaxxed the hell out of the tu…

Quantization is a massive efficiency gain for near negligible drop in quality. If the tradeoff is quantization for an 80 percent price drop I would take that any day of the week.

Re: OpenAI dropped the price of o3 by 80%

#24

...how? I'd understand a 20-30% price drop from infra improvements for a model as-is, but 80%? I wonder if "we quantized it lol" would classify as false advertising for modern LLMs.

Deepseek made a few major innovations allowing them to achieve major compute efficiency and then published them. My guess is that OpenAI just implemented these themselves.

Re: OpenAI dropped the price of o3 by 80%

#25

how do we know it's not a quantized version of o3? what's stopping these firms from announcing the full model to perform well on the benchmarks and then gradually quantizing it (first at Q8 so no one notices, then Q6, then Q4, ...). I have a suspicion that's how they were able to get gpt-4-turbo so fast. In practice, I found it inferior to the original GPT-4 but the company probably benchmaxxed the hell out of the tu…

Quantization is a massive efficiency gain for near negligible drop in quality. If the tradeoff is quantization for an 80 percent price drop I would take that any day of the week.

> for near negligible drop in quality

Hmm, that's evidently and anecdotally wrong:

https://github.com/ggml-org/llama.cpp/discussions/4110

Re: OpenAI dropped the price of o3 by 80%

#26

how do we know it's not a quantized version of o3? what's stopping these firms from announcing the full model to perform well on the benchmarks and then gradually quantizing it (first at Q8 so no one notices, then Q6, then Q4, ...). I have a suspicion that's how they were able to get gpt-4-turbo so fast. In practice, I found it inferior to the original GPT-4 but the company probably benchmaxxed the hell out of the tu…

I swear every time a new model is released it's great at first but then performance gets worse over time. I figured they were fine-tuning it to get rid of bad output which also nerfed the really good output. Now I'm wondering if they were quantizing it.

[flagged]

Re: OpenAI dropped the price of o3 by 80%

#27

It's going to be a race to the bottom, they have no moat.

Especially now that they are second in the race (behind Anthropic) and lot of free-to-download and free-to-use models are now starting to be viable competitors.

Once new MacBooks and iPhones have enough memory onboard this is going to be a disaster for OpenAI and other providers.

Re: OpenAI dropped the price of o3 by 80%

#28

You know. because LLMs can only be built by corporations... but because they're so easy to build, I see the price going down massively thanks to competition. Consumers benefit because all the companies are trying to out run each other.

And then they all go out of business, since models cost a fortune to build, and their fan club is left staring at their computers trying to remember how to do anything without getting it served on a silver plate.

Re: OpenAI dropped the price of o3 by 80%

#29

always seemed to me that efficient caching strategies could greatly reduce costs… wonder if they cooked up something new

How are LLMs cached? Every prompt would be different so it's not clear how that would work. Unless you're talking about caching the model weights...

Re: OpenAI dropped the price of o3 by 80%

#30
post #26

Earlier quoted context omitted.

I swear every time a new model is released it's great at first but then performance gets worse over time. I figured they were fine-tuning it to get rid of bad output which also nerfed the really good output. Now I'm wondering if they were quantizing it.

[flagged]

It's still a very competitive marketplace
Post reply on HN