Live data from Hacker News

DeepSeek API Pricing Update

api-docs.deepseek.com

31–40 of 201 posts

Re: DeepSeek API Pricing Update

#31
post #26

Earlier quoted context omitted.

How could that possibly work? Deepseek was undercutting every other provider by an order of magnitude on cached tokens. Do they just set a super low caching time and hope that drops effective cache rates low enough? Do all other providers somehow overcharge by that much? Are they just going to sell it as a loss leader?

> Do all other providers somehow overcharge by that much? This, I think. Cached inputs have an opportunity cost (keeping the KV cache until use) but a hit is basically free. “Basically” - if the cache is offloaded to system RAM or NVMe there’s some scheduling overhead. From a consumer viewpoint a more interesting metric than the raw costs is cached cost * hitrate + input cost * (1 - hitrate) from a personal standpoin…

System RAM and/or NVMe storage still has a real cost. And swapping out the context between VRAM and system RAM / NVMe still consumes bandwidth.

I don't have a clue on what the real cost to inference providers comes out to, but it seems really weird that there would be such a big gap, in what should be a pretty competitive market.

Re: DeepSeek API Pricing Update

#32
post #26

Earlier quoted context omitted.

How could that possibly work? Deepseek was undercutting every other provider by an order of magnitude on cached tokens. Do they just set a super low caching time and hope that drops effective cache rates low enough? Do all other providers somehow overcharge by that much? Are they just going to sell it as a loss leader?

> Do all other providers somehow overcharge by that much? This, I think. Cached inputs have an opportunity cost (keeping the KV cache until use) but a hit is basically free. “Basically” - if the cache is offloaded to system RAM or NVMe there’s some scheduling overhead. From a consumer viewpoint a more interesting metric than the raw costs is cached cost * hitrate + input cost * (1 - hitrate) from a personal standpoin…

> if OpenRouter is blindly dispatching your requests

This can somewhat be the case, depending on your config. I updated mine to make DeepSeek high priority because I was having a lot of cache misses and reliability issues with the default (cheapest (at face value)) providers, and cost was actually higher overall than anticipated. Was smooth sailing from then; might have to tweak things again now pricing has changed though.

Re: DeepSeek API Pricing Update

#33
post #25

Old DeepSeek Flash 0731 prices have been independently reproduced.[1] The issue is DeepSeek being inundated and not having capacity to serve the demand, hence the price increases to significantly dampen demand. Never mind international demand either--just think about the magnitude of Chinese domestic demand. Prices for anything related to AI or computing in general (mobile phones, cloud data centre hosting, etc) will…

Reproduced on the CUDA stack right?

Let's say DeepSeek is being forced to use the CANN stack, and the new pricing reflects the cost when 100% of inference is done with Huawei chips. Then, I suppose we can infer that:

* CANN stack is 1.5x~2.3x less efficient in compute

* CANN stack has 6x lower inter-connect capacity

> computer chips once again become a commodity

Ascend 950 is going for $7k to $9k with mediocre looking specs. $16k for RTX Pro 6000, $6k for RTX Pro 5000. This is not looking good.

Re: DeepSeek API Pricing Update

#34

According to their post: (Input / Output / Cache Read, [$/M]) DeepSeek-V4-Flash: Prev: 0.14 / 0.28 / 0.0028 Off-Peak: 0.22 (1.6x) / 0.66 (2.4x) / 0.007 (2.5x) Peak: 0.44 (3.1x) / 1.32 (4.7x) / 0.014 (5.0x) DeepSeek-V4-Pro: Prev: 0.435 / 0.87 / 0.003625 Off-Peak: 0.66 (1.5x) / 1.98 (2.3x) / 0.022 (6.1x) Peak: 1.32 (3.0x) / 3.96 (4.6x) / 0.044 (12.1x) gpt-5.6-luna: $0.20 / $1.20 / $0.02 / $0.25 (In / Out / Cache Read /…

> EDIT: formatting Keep at it, I believe in you.

My apologies to any mobile users, but for the desktop folk:

  Provider, Model    Billing       Input           Output          Cache read       Cache write
  DeepSeek                                                                            
    V4-Flash         Old           $0.1400         $0.2800         $0.0028          -      
    V4-Flash         New Off-Peak  $0.2200 (1.6x)  $0.6600 (2.4x)  $0.0070  (2.5x)  -      
    V4-Flash         New Peak      $0.4400 (3.1x)  $1.3200 (4.7x)  $0.0140  (5.0x)  -      
    V4-Pro           Old           $0.4350         $0.8700         $0.0036          -      
    V4-Pro           New Off-Peak  $0.6600 (1.5x)  $1.9800 (2.3x)  $0.0220  (6.1x)  -      
    V4-Pro           New Peak      $1.3200 (3.0x)  $3.9600 (4.6x)  $0.0440 (12.1x)  -      
                                                                                      
  OpenAI                                                                              
    GPT-5.6 Sol                    $5.0000         $30.000         $0.5000          $6.2500
    GPT-5.6 Terra                  $2.0000         $12.000         $0.2000          $2.5000
    GPT-5.6 Luna                   $0.2000         $1.2000         $0.0200          $0.2500
                                                                                      
  Anthropic                                                                           
    Claude Fable 5                 $10.000         $50.000         $1.0000          $12.500
    Claude Opus 5                  $5.0000         $25.000         $0.5000          $6.2500
    Claude Sonnet 5                $2.0000         $10.000         $0.2000          $2.5000
                                                                                      
  Moonshot                                                                            
    Kimi K3                        $3.0000         $15.000         $0.3000          -      
                                                                                      
  Z.AI                                                                                
    GLM 5.2                        $1.4000         $4.4000         $0.2600          -

Re: DeepSeek API Pricing Update

#35
post #25

Old DeepSeek Flash 0731 prices have been independently reproduced.[1] The issue is DeepSeek being inundated and not having capacity to serve the demand, hence the price increases to significantly dampen demand. Never mind international demand either--just think about the magnitude of Chinese domestic demand. Prices for anything related to AI or computing in general (mobile phones, cloud data centre hosting, etc) will…

[deleted]

Re: DeepSeek API Pricing Update

#36
I pay for Google AI Pro (Bought a year in advance) and Gemini is so bad, I burned through 75% of my five hour allowance trying to get it to fix something.

I pasted the same prompt into OpenCode, set to Deepseek v4 flash free and did it first try.

I'm was going to purchase Opencode GO to try it, but seems my timing is really bad :( hope it doesn't go up too much in Opencode or they find other providers. Bad timing!

Re: DeepSeek API Pricing Update

#37
post #31

Earlier quoted context omitted.

> Do all other providers somehow overcharge by that much? This, I think. Cached inputs have an opportunity cost (keeping the KV cache until use) but a hit is basically free. “Basically” - if the cache is offloaded to system RAM or NVMe there’s some scheduling overhead. From a consumer viewpoint a more interesting metric than the raw costs is cached cost * hitrate + input cost * (1 - hitrate) from a personal standpoin…

System RAM and/or NVMe storage still has a real cost. And swapping out the context between VRAM and system RAM / NVMe still consumes bandwidth. I don't have a clue on what the real cost to inference providers comes out to, but it seems really weird that there would be such a big gap, in what should be a pretty competitive market.

[deleted]

Re: DeepSeek API Pricing Update

#38
So OpenAI cut Luna's price by 5x, DeepSeek increased price by 5x!

If I'm reading the benchmarks right, they now went from being much cheaper than Luna (but twice as slow), to being roughly same price (but twice as slow).

So all else being equal, where I would previously have used DeepSeek, I can just use Luna, and get the same result twice as fast?

(Yeah I know benchmarks are mostly nonsense, but the ones measuring time are real, and it's the most precious resource.)

Re: DeepSeek API Pricing Update

#39

Earlier quoted context omitted.

> Do all other providers somehow overcharge by that much? This, I think. Cached inputs have an opportunity cost (keeping the KV cache until use) but a hit is basically free. “Basically” - if the cache is offloaded to system RAM or NVMe there’s some scheduling overhead. From a consumer viewpoint a more interesting metric than the raw costs is cached cost * hitrate + input cost * (1 - hitrate) from a personal standpoin…

> if OpenRouter is blindly dispatching your requests This can somewhat be the case, depending on your config. I updated mine to make DeepSeek high priority because I was having a lot of cache misses and reliability issues with the default (cheapest (at face value)) providers, and cost was actually higher overall than anticipated. Was smooth sailing from then; might have to tweak things again now pricing has changed t…

I was having the same issue. Did not find any provider that had a cache hit rate anywhere close to DeepSeeks own API.

Not sure if this was an OpenRouter issue or with the other inference providers.

Re: DeepSeek API Pricing Update

#40

About 3x increase. Luna is now a much better deal. Hope they don't increase their prices in response.

Luna vs v4 Flash, sure. But DeepSeek v4 Pro is a far more capable model and still cheaper than anything that it competes with, from what I can see.

Luna is multimodal.
Post reply on HN