If they can keep up this cadence of Flash leap-frogging the previous Pro, we're in for a good time
Anthropic/OpenAI might step up their anti-distillation defences though.
DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
61–70 of 214 posts
Re: DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
#62Sounds nice! But, the web ui chat version of flash has very poor language following abilities in my experience: You may ask it something in English, and get a thinking chain in Chinese with an answer in Chinese, or an English thinking chain and an English answer. Using the retry button on the same question has a 50/50 chance of any of those results. Sometimes, asking something in English, but where information are mo…
It's not just web chat, V4 Flash 7/31 suffers from a lot of pathological behavior in coding harnesses as well, e.g. infinite loops, hallucinations, premature termination, and invalid tool calls.
Re: DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
#63I've been watching a bunch of bycloud on YouTube recently, and although he's done a great job reviewing papers from the big AI labs, I feel like I'm missing something - how have all the labs seemingly made a model that's cheaper, faster AND has better performance? Historically `flash` variants (like codex spark as well) have been faster but perform worse
Re: DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
#64Re: DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
#65But what no one mentions is that the price is going from a starting point of $0.16 to $0.60, so basically they're charging nearly four times as much.
What are you talking about? Current flash prices are 0.66 for output, this is dropping it to 0.60.
We will adjust the pricing for the Flash series effective from 12:00 Beijing Time on September 10, 2026. During off-peak hours, the unit price will be $0.003 for input cache hits, $0.15 for input cache misses, and $0.6 for output. Peak-hour prices will be double the off-peak rates. Please plan your usage accordingly.
------------------------------------------------------- Hoje em sites como openrouter o valor é de $0.16 output .
Re: DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
#66https://www.geeky-gadgets.com/deepseek-v4-1-flash-review/
I hope some of those speed increases will make it to production.
Re: DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
#67Please don't do this kind of thing. If a user has validated a workflow on V4 Pro, they might not want to suddenly start testing it in production on V4.1 Flash. Instead, keep V4 Pro around but deprecated for a defined period of time, then remove it.
At least as open weights models, it's possible to use something like Together.ai or OpenRouter to run the V4 Pro model as long as other providers keep it up.
Re: DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
#68Re: DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
#69Earlier quoted context omitted.
The new meta model is fast and very cheap as well, and when used through OpenCode you get quite a lot of free tokens. But meta is also THE surveillance company, so probably also not a good choice in your case.
If you work on open source projects, I don't care about surveilance. It's right there on Github with full history anyways.
Re: DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
#70Since a few months, I almost exclusively use the Chinese "flash" models for my needs. They are a joy and they cost pennies per answer. Great job.
My mental bias always kept me away from Chinese models. Because i know that china is a surveillance state and all the things we know about CCP. But after what we learned about OpenAI and how they most likely used user data to basically cheat in an open competition i think it does not matter which AI provider you use all of them will own your data and all of them can spy on you. So I am willing to switch to Chinese mo…