Live data from Hacker News

DeepSeek V4 Flash 0731

arcprize.org

271–280 of 478 posts

Re: DeepSeek V4 Flash 0731

#271
post #252

One of the best things about this version is that it is trained in the codex harness. It feels just as good as OpenAI models in using codex tools, but extremely cheap and with 1M context

Do you have any sense how using it with codex compares to OpenCode? It’s always a bit tricky picking the right harness (when you have options). Sometimes the differences are subtle but meaningful. But who has the time to run everything twice and compare all the time!

I don't have experience with opencode, so I couldn't tell you.

Codex is really good in my experience, especially due to its native sandboxing. Deepseek seems really well versed in its tools, including update_plan and knowing when to request sandbox escalation.

Re: DeepSeek V4 Flash 0731

#272
post #247

I've been using it extensively since the release and the best summary I can give is that it's good enough to use it for (almost) everything and cheap enough that the cost are irrelevant. I'm running it in Oh My Pi with a second instance running as "advisor" and even with 5-6 active sessions (effectively 12 streams) I'm struggling to spend more than 5 bucks per day. OpenCode Go even has double limits temporarily so fo…

Are you using it via the official DeepSeek API, or via a different model provider? If the former, it's worth noting that their cache read prices are one tenth that of every other provider ($0.0028/M vs $0.028/M), so folks who want to use a sovereign inference provider with a zero data retention policy likely won't see anywhere close to the same value.

Worth mentioning also that DeepSeek is the only provider in OpenRouter that was disabled-by-default until I enabled a setting: Allow paid endpoints that train on request data.

Re: DeepSeek V4 Flash 0731

#273

Earlier quoted context omitted.

That doesn’t make sense. It’s not like SOTA models are error free, yet we still use them. You use Fable 5 right? If that’s good enough for you now, why wouldn’t a Chinese model that’s as good as Fable 5 but at 10% the cost be good enough in 6 months?

I think we put up with Fable's occasional hiccups because there's nothing better at the moment. I use Claude Code semi-heavily for my small business, and the $100/mo I pay for that is a rounding error compared to the value it provides. If I can avoid spending an hour or two "massaging" the output from a lower-end model once, or it avoids introducing one load-bearing (sorry, couldn't resist) bug, then that's the entir…

The question low cost models will create: Why would you massage output?

Fable 5 is still going to mess things up at any sufficient complexity. The advantage of low cost models with "good enough" intelligence is they can recursively correct. Why? Because it is cheap. Proper requirements and tests and subagents take away increasing amounts of work, at a cost that is not prohibitive.

If you are reviewing code manually you might consider Fable 5 a worse option. As it articulates itself with higher confidence and you already know it is capable, you are may be more likely to miss a mistake. You know to be on guard with a junior engineer. Reviewing a senior who suddenly makes some weird stochastic mistake can be a lot harder. It would be like if the smartest human engineer you knew was capable of some random brainfart in the middle of their massive diff. Imo, much harder to deal with.

Of course, we should keep in mind Fable 5 is only expensive today. It will be cheaper in the future. Autonomous, recursive prompting and improvement is the clear end state. Especially for entities that will always have the budget for that at the SOTA frontier.

Re: DeepSeek V4 Flash 0731

#275
post #170

Earlier quoted context omitted.

It is true. I don't care about having infinite frontier-level intelligence, and I don't care if Fable can one-shot frobnicate a klaxelzorp with a benchmark performance of 97%. I doubt most people do, in fact. I just want something that meets the baseline level of intelligence needed to be a really, really good pair programming agent. It shouldn't have any silly dealbreaker issues involving laziness or hallucinations,…

I wonder when we crossed the "99 percentile of intelligence for 99% of the usecases" threshold. At this point, the gains seem to be right at the very edge of bleeding edge for narrow and specialized use cases, and wonder if it'll be a sort of diminishing return from here on.

Probably the best counterexample is the games they are able to design. It's still mostly AI slop, few would want to play.

Re: DeepSeek V4 Flash 0731

#276

Earlier quoted context omitted.

That doesn’t make sense. It’s not like SOTA models are error free, yet we still use them. You use Fable 5 right? If that’s good enough for you now, why wouldn’t a Chinese model that’s as good as Fable 5 but at 10% the cost be good enough in 6 months?

I think we put up with Fable's occasional hiccups because there's nothing better at the moment. I use Claude Code semi-heavily for my small business, and the $100/mo I pay for that is a rounding error compared to the value it provides. If I can avoid spending an hour or two "massaging" the output from a lower-end model once, or it avoids introducing one load-bearing (sorry, couldn't resist) bug, then that's the entir…

> I think we put up with Fable's occasional hiccups because there's nothing better at the moment.

Which was an argument for using every less powerful model since the moment they got useful, right?

When was that? Opus 4.5 maybe? Let's say Opus 4.5 for the sake of the argument. So back then we were like "DeepSeek is not good enough, I need Opus 4.5". Now DeepSeek is better than Opus 4.5. So if Opus 4.5 was good enough back then, DeepSeek is better than that now.

Sure, it's always nicer to have a slightly better model. But the price difference starts mattering a lot more when all the models are already sufficiently good.

Re: DeepSeek V4 Flash 0731

#277
post #162

I've been using it extensively since the release and the best summary I can give is that it's good enough to use it for (almost) everything and cheap enough that the cost are irrelevant. I'm running it in Oh My Pi with a second instance running as "advisor" and even with 5-6 active sessions (effectively 12 streams) I'm struggling to spend more than 5 bucks per day. OpenCode Go even has double limits temporarily so fo…

If what you're saying is true and accurate, then US-based AI labs are in big trouble. The only saving grace might be some sort of a 'national security' proclamation banning the use of state-of-the-art Chinese (and non-US) models across US federal and state governments and large enterprises (especially ones with federal government contracts), but even still, US AI labs will probably lose out massively on international…

> only saving grace might be some sort of a 'national security' proclamation banning the use of state-of-the-art Chinese (and non-US) models

In what kind of sad and failed dystopia is this a "saving grace"? For whom?

Re: DeepSeek V4 Flash 0731

#278

This latest DeepSeek is almost at the "too cheap to meter" level. That's going to be a larger unlock than models like Fable/Mythos that are way too expensive to justify, IMO. What secret sauce do they have?

> What secret sauce do they have?

Quant company usually squeezing every penny.

Re: DeepSeek V4 Flash 0731

#279
post #268
post #162

Earlier quoted context omitted.

If what you're saying is true and accurate, then US-based AI labs are in big trouble. The only saving grace might be some sort of a 'national security' proclamation banning the use of state-of-the-art Chinese (and non-US) models across US federal and state governments and large enterprises (especially ones with federal government contracts), but even still, US AI labs will probably lose out massively on international…

The bet isn't that people will be able to automatically reply on bugs and rack up API charges. The bet is on using AI to gain competitive advantage. You don't win the stock market or make the deadliest drone by switching to the cheap model

> You don't win the stock market or make the deadliest drone by switching to the cheap model

Really? How many times a small team has outperformed a much bigger one just because they were "doing it right"?

I have been in software companies where most software produced was bad. Not just the code, the overall design everywhere. So... bad engineers with the most expensive model, or great engineers with cheaper models?

Re: DeepSeek V4 Flash 0731

#280
post #258

Earlier quoted context omitted.

It's specific to the harness. Using dynamic context pruning, the budget cuts it off after a select amount of tokens and the budget message tells the model to use subgents to finish whatever it's thinking about

Nice. Does it use a summarization, or a hard cutoff?

llama.cpp uses a hard cutoff. The agent then does "something" that is specific to the agent's implementation and configuration. It might summarize and then "finish the thought" with a different model, and then resubmit the prompt to the llama.cpp API endpoint with .. prefilled. The primary model then infers the remainder of the reply.
Post reply on HN