Live data from Hacker News

DeepSeek V4 Flash 0731

arcprize.org

91–100 of 477 posts

Re: DeepSeek V4 Flash 0731

#91

Earlier quoted context omitted.

Dax (from Opencode) has tweeted that they can replicate or beat the price with rented GPUs. Deepseeks secret sauce is the incredibly cheap caching (magnitude cheaper than other providers). vLLM has recently released a similar approach. It's not as effective as what DeepSeek does but still an interesting development. I have no doubt that in due time other providers will match or perhaps even beat the current DeepSeek…

> Deepseeks secret sauce is the incredibly cheap caching (magnitude cheaper than other providers). Can anyone working at one of the main US labs (Google, OpenAI, Anthropic) comment on WTF they haven't even tried MLA - despite the obvious massive advantages? I know enough to know they aren't completely incompetent. So there must be a quite good reason. But it remains a mystery to me. DeepSeek's MLA is like almost 2 ye…

They already are?

There’s a measurable performance tradeoff versus gqa so there’s reluctance.

For the most part though the new deepseek v4 tech is hca and mhc and people are still catching on like with moe and rl. Wait for 6 12 months, minimum time for next pre train.

Re: DeepSeek V4 Flash 0731

#92
post #2

results comparable to gpt 5.6 luna but cheaper promising!

Might not actually be that much cheaper, we don't know what margin OpenAI is charging on Luna API. Open models likely have much less margin.

Re: DeepSeek V4 Flash 0731

#93

I've been using it extensively since the release and the best summary I can give is that it's good enough to use it for (almost) everything and cheap enough that the cost are irrelevant. I'm running it in Oh My Pi with a second instance running as "advisor" and even with 5-6 active sessions (effectively 12 streams) I'm struggling to spend more than 5 bucks per day. OpenCode Go even has double limits temporarily so fo…

These posts have to be Chinese bots, these models are all trash. Used it via OpenCode for an hour, cost me one hour of my life. It is for anything complete trash.

Re: DeepSeek V4 Flash 0731

#94
post #15

Earlier quoted context omitted.

But DeepSeek now has a warning they’re going to sharply increase their API pricing sometime in the future.

Dax (from Opencode) has tweeted that they can replicate or beat the price with rented GPUs. Deepseeks secret sauce is the incredibly cheap caching (magnitude cheaper than other providers). vLLM has recently released a similar approach. It's not as effective as what DeepSeek does but still an interesting development. I have no doubt that in due time other providers will match or perhaps even beat the current DeepSeek…

I'll believe it when I see it. Their prices are still much higher than deepseek, especially the caching.

Re: DeepSeek V4 Flash 0731

#95
post #6
post #4

Kimi K3 was an interesting model only a month ago, and now we're looking at the same performance for 1/20th of the price. Wild how fast this is advancing.

Yeah either the benchmark isn't very useful anymore or V4 Flash is a really, really good model.

GPT 5.6 Luna is an extremely cheap and still very capable model.

A chinese model being in the same ballpark of capability at half the price sounds believable to me.

Re: DeepSeek V4 Flash 0731

#96
post #93

I've been using it extensively since the release and the best summary I can give is that it's good enough to use it for (almost) everything and cheap enough that the cost are irrelevant. I'm running it in Oh My Pi with a second instance running as "advisor" and even with 5-6 active sessions (effectively 12 streams) I'm struggling to spend more than 5 bucks per day. OpenCode Go even has double limits temporarily so fo…

These posts have to be Chinese bots, these models are all trash. Used it via OpenCode for an hour, cost me one hour of my life. It is for anything complete trash.

(1) you used opencode (2) what provider did you use. openrouter is trash because they shit up the model serving. no max effort and horrific cache utilization, on the order of 50-75%, absolutely garbage. beware

Re: DeepSeek V4 Flash 0731

#97
post #44

I strongly recommend trying this for programming tasks. It is strong (not Fable strong though) with a much better “persona” than Opus, and very different blindspots. If you flip between Claude and this you will find both catch the mistakes of the other before they get out of control. On balance I actually prefer DeepSeek for programming now, because of the way it talks.

[flagged]

Re: DeepSeek V4 Flash 0731

#100
post #88

I've been using it extensively since the release and the best summary I can give is that it's good enough to use it for (almost) everything and cheap enough that the cost are irrelevant. I'm running it in Oh My Pi with a second instance running as "advisor" and even with 5-6 active sessions (effectively 12 streams) I'm struggling to spend more than 5 bucks per day. OpenCode Go even has double limits temporarily so fo…

> it's good enough to use it for (almost) everything which in your case is?

> which in your case is?

oh, they're mad.

Post reply on HN