Live data from Hacker News

DeepSeek V4 Flash 0731

arcprize.org

461–470 of 474 posts

Re: DeepSeek V4 Flash 0731

#461
post #403

Earlier quoted context omitted.

> The Chinese models are not good enough for anything other than pair programming, which is just a very last-gen way of using agents. This matches my experience with DeepSeek V4 Pro at Max reasoning, the preview version of the model kept regularly messing things up. About 30-60% of additional time to fix the output was needed. On similar tasks, GLM 5.2 at Max reasoning screwed up maybe 20-30% of the time, while it st…

> This matches my experience with DeepSeek V4 Pro at Max reasoning, That was ages ago (in LLM release timelines). DeepSeek V4 Flash beats it now and a lot cheaper. > On similar tasks, GLM 5.2 at Max reasoning screwed up maybe 20-30% of the time, GLM 5.3 bridges this gap. > I'd say as Chinese models get better, whatever moat Anthropic and OpenAI have dissipates. Their moat, especially OpenAI is funding and hardware re…

> GLM 5.3 bridges this gap.

I’m sure the next models will only get better, when they’re released. Also super curious about what Moonshot will achieve and the full DeepSeek V4 Pro release!

> Their moat, especially OpenAI is funding and hardware resources. They gain train models 10x as large and also serve at large scale. That's it.

I’ve seen how much slower Kimi K3 can be and that part seems correct, their own GPU production still has ways to go and export restrictions definitely limit what they can do.

Not sure about the size part, if Kimi K3 achieves SOTA performance at 2.8T parameters, western models being >2x that size would be insanely bad in regards to efficiency. I bet they’re all within the same order of magnitude and below 10T and won’t really have a reason to go even that high for the foreseeable future.

As investors will start squeezing them for profitability, I suspect focusing more on efficiency will be commonplace.

Re: DeepSeek V4 Flash 0731

#462
post #2

results comparable to gpt 5.6 luna but cheaper promising!

Is it still cheaper than Luna if using an OpenAI subscription? My gut is no, but I have not done the math.

Some china proxies charge as little as 1 cent / Mio tokens for luna, which I use. As far as I know they bundle multiple codex subscriptions to get subsidized token pricings, so in codex it should be.

Re: DeepSeek V4 Flash 0731

#463
post #162

Earlier quoted context omitted.

If what you're saying is true and accurate, then US-based AI labs are in big trouble. The only saving grace might be some sort of a 'national security' proclamation banning the use of state-of-the-art Chinese (and non-US) models across US federal and state governments and large enterprises (especially ones with federal government contracts), but even still, US AI labs will probably lose out massively on international…

I think you're overlooking the fact that for long-horizon tasks, even small errors compound over time and can lead to catastrophic outcomes. For simple queries, we have reached the threshold since the beginning of the year, and models are good enough from every provider to make a meaningful difference between one another. (ChatGPT, Claude, Gemini, Grok, MuseSpark, Kimi, DeepSeek, GLM...) The real unlock will be, and…

Amazing how many inaccuracies you fit in there.

- You describe the breakdown in terms of time but it's more accurately a function of reasoning complexity.

- You seem to assume that no intermediate evaluation is possible.

- Often it is (e.g. the build breaks or tests start failing), allowing for course correction. There's definitely a cost to that but it can still be cost effective if the accuracy is "good enough" and the price difference significant.

- There are numerous tasks that don't require Fable or GPT5.6 level reasoning to improve efficiency by an order of magnitude.

Re: DeepSeek V4 Flash 0731

#464

Earlier quoted context omitted.

To put actual numbers on it, since using AI to start solving all kinds of bottlenecks/inefficiencies in our small business, we've seen monthly net profit go up by around $4,000 USD. These are semi-permanent fixes, and the tech is only partially deployed. I am the only one using it, and I only use it part time. We've just spun up our first Hermes agent, with direct API access to our main inventory system and that's ex…

"оur first Hermes agent, with direct API access to our main inventory system" – let me assure you that absolutely nothing can go wrong here, mate. /s

Did you mean to post this on Reddit instead of HN?

Re: DeepSeek V4 Flash 0731

#465
post #441

Earlier quoted context omitted.

I think HN doesn't really understand fixed costs, I spend a few hundred dollars per day on Fable and the costs are irrelevant compared to what we make.

I've never understood this line of thinking... Are you saying we shouldn't care about the future of affordability and access because at this moment we have seemingly endless access? Sounds extremely short sighted.

No, its that going from $2k/mo to $20/mo is not worth the capability loss in the slightest, because AI just isn't a big part of our fixed costs.

Re: DeepSeek V4 Flash 0731

#466

Earlier quoted context omitted.

I use Chinese models, even smaller local ones, for much more than pair programming. If we are talking about deepseek v4 flash, which is basically a frontier model, it is much more capable than the local models I run on my MacBook Pro. The only issue really is finding the right harness. I do have a way of correcting through redundancy, though. If you are just vibe coding, you need to use the most capable model you can…

> If you are just vibe coding, you need to use the most capable model you can find and even then it might not be good enough I mean, this proves my point. Better models enable you to get more done. With Fable, 80% of the time, I no longer have chat with an agent over the details of a PR. I give it an outcome and it gets done. This means I can work on much more with the limited time I have. And I don't see this ending…

If you don’t have a way to automatically check the results via redundancy, you need a really good model, since even 99.9% reliability is going to cause slot of headaches. If you do have a way to automatically check results, then you can use something that fits in your computer and was produced 3 years ago.

My point is you don’t need the best model if you just put in QA processes that can be done by models also. And if you don’t have that, the model is probably not going to be good enough.

Re: DeepSeek V4 Flash 0731

#467

I've been using it extensively since the release and the best summary I can give is that it's good enough to use it for (almost) everything and cheap enough that the cost are irrelevant. I'm running it in Oh My Pi with a second instance running as "advisor" and even with 5-6 active sessions (effectively 12 streams) I'm struggling to spend more than 5 bucks per day. OpenCode Go even has double limits temporarily so fo…

Hard to make big predictions, but it sure looks like at least this level of capability is going to be available in the open and relatively cheap to run.

The 'floor' has gone up: today's model a bit behind SOTA is like model releases that were blowing people's minds a few months ago. Compared to, say, DS R1, this is far out stuff.

This, Luna, and (if it's good in practice) Laguna S are also fast and light not just cheap. And, as happened before, DeepSeek's first but other open model makers likely follow.

And a small, fast model taking small steps is...fun? More like working with code.

Re: DeepSeek V4 Flash 0731

#468

Earlier quoted context omitted.

Those are not at the same price. Opecode’s 60 USD of deepseek usage is charged at much higher rates than what deepseek themselves charge at.

> Opecode’s 60 USD of deepseek usage is charged at much higher rates than what deepseek themselves charge at Per the open code zen pricing page[1], it appears that the token prices are the same, but their cache is 10x more expensive? [1]: https://opencode.ai/docs/zen/#pricing

Flash Opencode 0.14/0.28/0.028

Deepseek 0.14/0.28/0.0028

Pro Opencode 1.74/3.48/0.145

Deepseek 0.44/0.87/0.0036

For flash the input/output is the same, but the cache difference is big, you're paying 10x on >95% of your tokens.

For Pro, it's even worse, input/output is 4x and cache is 40x. The price different is really brutal. Yes you will still come out ahead by spending your first 10$/month on opencode go, but you will be saving a lot less than initially appears from their (60 USD for 10 USD pitch).

[0]: Deepseek: https://api-docs.deepseek.com/quick_start/pricing/ [1]: Opencode: https://opencode.ai/docs/zen/#pricing

Re: DeepSeek V4 Flash 0731

#469

Earlier quoted context omitted.

This stuff is just obvious to anyone working with the latest models. Fully autonomous agents are a game changer. Having to pair program with one is indeed "last gen", I haven't done that for a month and I won't ever be doing that again in my life, outside of personal projects.

> This stuff is just obvious to anyone working with the latest models. Fully autonomous agents are a game changer. This kind of takes makes me cringe. Why don't you go back to LinkedIn? I have no idea what you mean by “pair program with an agent”, but Opus have been able of autonomous coding since last November, and with any half-decent harness even local Qwen3.5 was able to do so 6 months ago. Fable is a stronger mo…

> This kind of takes makes me cringe. Why don't you go back to LinkedIn?

This is not appropriate for HN. Please review the guidelines: https://news.ycombinator.com/newsguidelines.html

Re: DeepSeek V4 Flash 0731

#470
post #403

Earlier quoted context omitted.

> This matches my experience with DeepSeek V4 Pro at Max reasoning, That was ages ago (in LLM release timelines). DeepSeek V4 Flash beats it now and a lot cheaper. > On similar tasks, GLM 5.2 at Max reasoning screwed up maybe 20-30% of the time, GLM 5.3 bridges this gap. > I'd say as Chinese models get better, whatever moat Anthropic and OpenAI have dissipates. Their moat, especially OpenAI is funding and hardware re…

> GLM 5.3 bridges this gap. I’m sure the next models will only get better, when they’re released. Also super curious about what Moonshot will achieve and the full DeepSeek V4 Pro release! > Their moat, especially OpenAI is funding and hardware resources. They gain train models 10x as large and also serve at large scale. That's it. I’ve seen how much slower Kimi K3 can be and that part seems correct, their own GPU pro…

> Not sure about the size part, if Kimi K3 achieves SOTA performance at 2.8T parameters, western models being >2x that size would be insanely bad in regards to efficiency. I bet they’re all within the same order of magnitude and below 10T and won’t really have a reason to go even that high for the foreseeable future.

They're a lot larger e.g. Fable. It is insanely bad. Do you know how much more resources "Western" companies have? Most in China don't have random GPUs to "play with" like every "frontier lab" employee does.

> As investors will start squeezing them for profitability, I suspect focusing more on efficiency will be commonplace.

They're born lucky though. Efficiency is "free". The next generation hardware e.g. Nvidia claims Blackwell -> Rubin is 10x efficiency (verified by Neoclouds apparently).

Post reply on HN