Live data from Hacker News

DeepSeek launching v4.1 flash cheaper and more capable than v4 pro

news.ycombinator.com

61–70 of 221 posts

Re: DeepSeek launching v4.1 flash cheaper and more capable than v4 pro

#61

If they can keep up this cadence of Flash leap-frogging the previous Pro, we're in for a good time

Anthropic/OpenAI might step up their anti-distillation defences though.

They might need to step up their product offerings and offer cheaper.

Re: DeepSeek launching v4.1 flash cheaper and more capable than v4 pro

#62
post #20

Sounds nice! But, the web ui chat version of flash has very poor language following abilities in my experience: You may ask it something in English, and get a thinking chain in Chinese with an answer in Chinese, or an English thinking chain and an English answer. Using the retry button on the same question has a 50/50 chance of any of those results. Sometimes, asking something in English, but where information are mo…

It's not just web chat, V4 Flash 7/31 suffers from a lot of pathological behavior in coding harnesses as well, e.g. infinite loops, hallucinations, premature termination, and invalid tool calls.

FWIW, I haven’t experienced any of that using V4 Flash via DeepSeek in omp. What’s your coding harness and inference provider?

Re: DeepSeek launching v4.1 flash cheaper and more capable than v4 pro

#63
post #56

I've been watching a bunch of bycloud on YouTube recently, and although he's done a great job reviewing papers from the big AI labs, I feel like I'm missing something - how have all the labs seemingly made a model that's cheaper, faster AND has better performance? Historically `flash` variants (like codex spark as well) have been faster but perform worse

[dead]

Re: DeepSeek launching v4.1 flash cheaper and more capable than v4 pro

#65
post #32
post #28

But what no one mentions is that the price is going from a starting point of $0.16 to $0.60, so basically they're charging nearly four times as much.

What are you talking about? Current flash prices are 0.66 for output, this is dropping it to 0.60.

This is the notice from DeepSeek regarding their API:

We will adjust the pricing for the Flash series effective from 12:00 Beijing Time on September 10, 2026. During off-peak hours, the unit price will be $0.003 for input cache hits, $0.15 for input cache misses, and $0.6 for output. Peak-hour prices will be double the off-peak rates. Please plan your usage accordingly.

------------------------------------------------------- Hoje em sites como openrouter o valor é de $0.16 output .

Re: DeepSeek launching v4.1 flash cheaper and more capable than v4 pro

#67
>In keeping with our commitment to user responsibility, following the official launch of V4.1 Flash and prior to the release of V4.1 Pro, all requests to the Pro model will be routed to V4.1 Flash and billed at Flash's price

Please don't do this kind of thing. If a user has validated a workflow on V4 Pro, they might not want to suddenly start testing it in production on V4.1 Flash. Instead, keep V4 Pro around but deprecated for a defined period of time, then remove it.

At least as open weights models, it's possible to use something like Together.ai or OpenRouter to run the V4 Pro model as long as other providers keep it up.

Re: DeepSeek launching v4.1 flash cheaper and more capable than v4 pro

#69
post #35

Earlier quoted context omitted.

The new meta model is fast and very cheap as well, and when used through OpenCode you get quite a lot of free tokens. But meta is also THE surveillance company, so probably also not a good choice in your case.

If you work on open source projects, I don't care about surveilance. It's right there on Github with full history anyways.

Is it? You're still putting a lot of thought and guidance into the agent's harness, the final code is just a tiny bit of that. It's like giving a junior developer final code vs explaining the whys and nuance. Which I'm not sure I want to give Meta

Re: DeepSeek launching v4.1 flash cheaper and more capable than v4 pro

#70
post #15

Since a few months, I almost exclusively use the Chinese "flash" models for my needs. They are a joy and they cost pennies per answer. Great job.

My mental bias always kept me away from Chinese models. Because i know that china is a surveillance state and all the things we know about CCP. But after what we learned about OpenAI and how they most likely used user data to basically cheat in an open competition i think it does not matter which AI provider you use all of them will own your data and all of them can spy on you. So I am willing to switch to Chinese mo…

Even you think both cheat, you can't possibly think they both cheat the same amount.
Post reply on HN