Looking at the caching price of deepseek compared to its competitors, does it have a secret sauce or is it just subsidizing?
DeepSeek V4 Flash 0731
181–190 of 481 posts
Re: DeepSeek V4 Flash 0731
#182Does no thinking emissions for context saving.
Re: DeepSeek V4 Flash 0731
#183From here on, it's going to become all about harnesses that best situate and organize swarm intelligence at scale.
Re: DeepSeek V4 Flash 0731
#184Earlier quoted context omitted.
Others have said similar but I disagree, I'm spending $200/m and I can easily burn through my weekly quota with a few overnight goals using 5.6 medium.
And what do you do with all that?
I can pretty easily burn through my weekly quota over several agent coding hours with minimal supervision when tasked with some pretty large but well-planned refactors.
Re: DeepSeek V4 Flash 0731
#185Perhaps it might be interesting: a latent thinking version is here https://huggingface.co/nmitchko/DeepSeek-V4-Flash-0731-Laten... Does no thinking emissions for context saving.
Re: DeepSeek V4 Flash 0731
#186Note this is the 07/31 release of DSv4 flash and not the "preview" that they put out a couple months or so ago. I've been running this model locally for a week, and the preview version before that. This updated one feels like a whole tier up. It's very capable for debugging and analyzing documents/data I upload. The killer feature, IMO, is the speed. On 2x RTX Pro 6000 Blackwell, its ~8k tok/s prefill and ~250 tok/s…
What quantization level is that? Because official endpoints are slow .
I wouldn't call 80 t/s slow.
Re: DeepSeek V4 Flash 0731
#187Earlier quoted context omitted.
And now nobody seems interested in it because the price hasn't gone down it's still $3/$15 for all providers on openrouter because of some Kimi license https://openrouter.ai/moonshotai/kimi-k3#providers
Morph has it for a slight discount, apparently. Uptime looks crap, though.
So we won't see any price decrease unless Kimi changes the license of K3
Re: DeepSeek V4 Flash 0731
#188Note this is the 07/31 release of DSv4 flash and not the "preview" that they put out a couple months or so ago. I've been running this model locally for a week, and the preview version before that. This updated one feels like a whole tier up. It's very capable for debugging and analyzing documents/data I upload. The killer feature, IMO, is the speed. On 2x RTX Pro 6000 Blackwell, its ~8k tok/s prefill and ~250 tok/s…
What quantization level is that? Because official endpoints are slow .
Re: DeepSeek V4 Flash 0731
#189I've been using it extensively since the release and the best summary I can give is that it's good enough to use it for (almost) everything and cheap enough that the cost are irrelevant. I'm running it in Oh My Pi with a second instance running as "advisor" and even with 5-6 active sessions (effectively 12 streams) I'm struggling to spend more than 5 bucks per day. OpenCode Go even has double limits temporarily so fo…
If what you're saying is true and accurate, then US-based AI labs are in big trouble. The only saving grace might be some sort of a 'national security' proclamation banning the use of state-of-the-art Chinese (and non-US) models across US federal and state governments and large enterprises (especially ones with federal government contracts), but even still, US AI labs will probably lose out massively on international…
Re: DeepSeek V4 Flash 0731
#190I've been using it extensively since the release and the best summary I can give is that it's good enough to use it for (almost) everything and cheap enough that the cost are irrelevant. I'm running it in Oh My Pi with a second instance running as "advisor" and even with 5-6 active sessions (effectively 12 streams) I'm struggling to spend more than 5 bucks per day. OpenCode Go even has double limits temporarily so fo…