Live data from Hacker News

DeepSeek V4 Flash 0731

arcprize.org

401–410 of 481 posts

Re: DeepSeek V4 Flash 0731

#401

Earlier quoted context omitted.

Except Fable won’t be costing $100 for enterprises that will be considering the Chinese models. If $100 Claud Max subscription works for you, then great. But you have to remember your pricing is subsidized by enterprises that pay hundreds of thousands of dollars each month, if not more, to Anthropic. For those companies, a Chinese model that can cut their AI spend from $1M/month to $200k suddenly seems attractive. An…

Didn’t deepseek recently announce prices will go up significantly? Right now the US dominates everyone else in actual chips in data centers. So even if deepseek etc tries to undercut, they’re very capacity limited.

> Didn’t deepseek recently announce prices will go up significantly?

They also previously said prices will go down significantly once they get a hold of the upcoming Huawei chips (later this year).

Prices are going up just because they can. It can easily come back down. They aren't strained by some IPO / VCs requiring them to 1000x their earnings.

Re: DeepSeek V4 Flash 0731

#402

For the last 3 months I've been using V4 Flash Free with Hermes through Opencode Zen both personally and at my company and I've been having a great experience so far. It's my go-to model for terminal work, managing my entire ubuntu server, Cloudpanel, managing static websites, doing SEO audits, network tests, DNS troubleshooting, e-mail deliverability troubleshooting... Furthermore, in my company we are using MCPs fo…

Sounds super simple stuff compared to what I use Fable and Opus 5 for.

Re: DeepSeek V4 Flash 0731

#403

Earlier quoted context omitted.

I work for a FAANG, and have my own personal projects for which I use the Chinese models. and the big models do indeed save/make us a lot of money. The Chinese models are not good enough for anything other than pair programming, which is just a very last-gen way of using agents. And when the big US models get better we will move with them. Until we stop seeing returns there is no "good enough", I don't know why this…

> The Chinese models are not good enough for anything other than pair programming, which is just a very last-gen way of using agents. This matches my experience with DeepSeek V4 Pro at Max reasoning, the preview version of the model kept regularly messing things up. About 30-60% of additional time to fix the output was needed. On similar tasks, GLM 5.2 at Max reasoning screwed up maybe 20-30% of the time, while it st…

> This matches my experience with DeepSeek V4 Pro at Max reasoning,

That was ages ago (in LLM release timelines). DeepSeek V4 Flash beats it now and a lot cheaper.

> On similar tasks, GLM 5.2 at Max reasoning screwed up maybe 20-30% of the time,

GLM 5.3 bridges this gap.

> I'd say as Chinese models get better, whatever moat Anthropic and OpenAI have dissipates.

Their moat, especially OpenAI is funding and hardware resources. They gain train models 10x as large and also serve at large scale. That's it.

Re: DeepSeek V4 Flash 0731

#404

Earlier quoted context omitted.

You're using it on low, that's why. There's a huge difference in performance from low to max effort.

I’m comparing same / similar settings between models. I can’t use high on one and low on others it’s not a fair test. Not sure why I was downvoted. But seems the downvoter is quick to downvote anything that doesn’t fit the narrative they’re looking for. I’m just reporting my findings.

Low, High and Max, obviously, can't be compared across models. They only mean the model is likely to spend less reasoning effort (~output tokens) with Low than High on the same, *single shot* task.

But even in this very post, you can see that Max was actually cheaper than High.

If you are using API, you should be comparing based on end-to-end cost or speed or whatever blend of those two matches your cost/time budget.

Re: DeepSeek V4 Flash 0731

#406
post #375

Earlier quoted context omitted.

It's significantly worse than Luna and quite a bit slower in some fairly involved tests I run.

That's fascinating, it's WAY better than luna ime. What sort of things are you testing it for?

I've been using this DeepSeek model the whole day today after building with 5.6 Luna extensively over the last week and I would disagree, at least for Rust + OpenGL.

DeepSeek just spend almost 2 hours trying to figure out why terrain textures were not working. It tried everything over and over again, it even had reference code for meshes on how to setup the rendering with materials, and it could just not do it.

I finally gave up and gave it to GPT-5.6 Luna instead, and figure out in a single prompt after 20 seconds, that the terrain mesh was being initialized with None in the material slot.

Other tasks it has managed to figure out at least, but it is significantly slower than GPT-5.6 Luna and it requires a lot more iterations.

(Both were set to high reasoning)

Re: DeepSeek V4 Flash 0731

#407

Perhaps it might be interesting: a latent thinking version is here https://huggingface.co/nmitchko/DeepSeek-V4-Flash-0731-Laten... Does no thinking emissions for context saving.

This is pretty interesting, I've never heard of this approach before - do you know if there is a research paper that covers how this was achieved?

It’s an adaptation of CoLaR, but my implementation is a little different: - Dedicated stop head to fire when latent thinking hits threshold - MTP support with training taking draft support as first class - different architectural layer 35 -> layer 42 writeback. So latents skip roughly 6.2 tokens of reasoning per token, then never make it to decoded output

Re: DeepSeek V4 Flash 0731

#408
post #366

Earlier quoted context omitted.

Probably the best counterexample is the games they are able to design. It's still mostly AI slop, few would want to play.

With those cheap models the idea is you're still in the loop anyway so the more expensive model is a waste of time and money. In this case, that means you're steering the game to look like you want not how AI wants.

Same with the expensive models. Fable can't one-shot a good new game. Game dev is still human-in-the-loop no matter what model you're using.

Re: DeepSeek V4 Flash 0731

#409

Earlier quoted context omitted.

"оur first Hermes agent, with direct API access to our main inventory system" – let me assure you that absolutely nothing can go wrong here, mate. /s

Don't you think that's a bit of an inane comment to make about a setup you know nothing about? Surely they are using read-only access.

We are obviously limiting access to be read-only for anything customer-facing (it will be able to put recommendations in the dashboard but not actually change things directly) but I think the point of GP's comment was literally just to get a reaction.

Re: DeepSeek V4 Flash 0731

#410

Earlier quoted context omitted.

I think we put up with Fable's occasional hiccups because there's nothing better at the moment. I use Claude Code semi-heavily for my small business, and the $100/mo I pay for that is a rounding error compared to the value it provides. If I can avoid spending an hour or two "massaging" the output from a lower-end model once, or it avoids introducing one load-bearing (sorry, couldn't resist) bug, then that's the entir…

The discussion started about cost-vs-usability, and then you brought in "but humans, the best LLMs around, cost much more" into this discussion to make it a not cost-vs-usability discussion. Do you not find it a bit disingenuous?

Not really, no. My point is that a $100/mo LLM subscription is generating $5000/mo in value, so even a negligible difference in performance wipes out the cost savings completely.

Same reason it makes sense to assign a team of humans that cost $100k/mo to a product that brings in $5M/mo, rather than one human with 5 Claude Max subs.

The cost is a rounding error.

Post reply on HN