Live data from Hacker News

DeepSeek V4 Flash 0731

arcprize.org

311–320 of 477 posts

Re: DeepSeek V4 Flash 0731

#312

I've been using it extensively since the release and the best summary I can give is that it's good enough to use it for (almost) everything and cheap enough that the cost are irrelevant. I'm running it in Oh My Pi with a second instance running as "advisor" and even with 5-6 active sessions (effectively 12 streams) I'm struggling to spend more than 5 bucks per day. OpenCode Go even has double limits temporarily so fo…

Hear hear. IQ tokens to cheap to meter upon us. So many things changed since last week. Now I've had Prime agent session grinding into its 20-th hour still not giving up. Been using opencode-go since Go sub appeared. What made a difference was deepseek-v4-flash and mimo-v2.5 showing. Very similar middling models ~300b so light on the gpu. 1M context and hybrid archs - so one can actually make use of that 1M (don't gr…

I would not recommend DSV4F (even 0731 edition) without an advisor. On its own it’s an absolute drunk intern in my experience, but with an advisor model watching like a hawk when it gets stuck in loops or goes down boneheaded rabbit holes, it’s fine (and very cheap). I’ve been using GLM-5.2 as my /advisor but might try just a second DSV4F instance.

Re: DeepSeek V4 Flash 0731

#313
post #276

Earlier quoted context omitted.

> I think we put up with Fable's occasional hiccups because there's nothing better at the moment. Which was an argument for using every less powerful model since the moment they got useful, right? When was that? Opus 4.5 maybe? Let's say Opus 4.5 for the sake of the argument. So back then we were like "DeepSeek is not good enough, I need Opus 4.5". Now DeepSeek is better than Opus 4.5. So if Opus 4.5 was good enough…

To put actual numbers on it, since using AI to start solving all kinds of bottlenecks/inefficiencies in our small business, we've seen monthly net profit go up by around $4,000 USD. These are semi-permanent fixes, and the tech is only partially deployed. I am the only one using it, and I only use it part time. We've just spun up our first Hermes agent, with direct API access to our main inventory system and that's ex…

"оur first Hermes agent, with direct API access to our main inventory system" – let me assure you that absolutely nothing can go wrong here, mate. /s

Re: DeepSeek V4 Flash 0731

#314
post #276

Earlier quoted context omitted.

> I think we put up with Fable's occasional hiccups because there's nothing better at the moment. Which was an argument for using every less powerful model since the moment they got useful, right? When was that? Opus 4.5 maybe? Let's say Opus 4.5 for the sake of the argument. So back then we were like "DeepSeek is not good enough, I need Opus 4.5". Now DeepSeek is better than Opus 4.5. So if Opus 4.5 was good enough…

To put actual numbers on it, since using AI to start solving all kinds of bottlenecks/inefficiencies in our small business, we've seen monthly net profit go up by around $4,000 USD. These are semi-permanent fixes, and the tech is only partially deployed. I am the only one using it, and I only use it part time. We've just spun up our first Hermes agent, with direct API access to our main inventory system and that's ex…

I think we're even starting to reach that saturation point now for a lot of people. In my industry (law) plenty of people have tried CoPilot once or twice, or tried ChatGPT a year ago, and as a result have basically dismissed AI as being useless. The setup required to be able to get it to do end to end tasks to your liking is also substantially more work than most people are willing to put in.

Re: DeepSeek V4 Flash 0731

#315
I'm not sure about all these benchmarks, I did some very simple tests (I have my own benchmarks https://upmaru.com/llm-tests) and these models fail, not sure if it's the inference provider or the model. They seem to be optimized for benchmarks more than real use cases. Do anything outside their distribution (even if it's not complex) they fail.

I Compared Deepseek V4 Flash 0731 (low) to Gemini 3.5 Flash Lite (minimal) and GPT 5.6 Luna (no reasoning) and Deepseek V4 Flash 0731 gets it wrong alot, where as Gemini and 5.6 Luna just gets it done.

Re: DeepSeek V4 Flash 0731

#317

Earlier quoted context omitted.

Just trying to understand, https://opencode.ai/docs/go/#privacy currently says DeepSeek V4 Flash has 0 days data retention. > DeepSeek V4 Flash: ZDR agreement is renewed monthly. The current agreement is valid through August 31, 2026. Is there other info I should be aware of w.r.t data retention with opencode go? It's hosted in China, so other middlemen may be active (I doubt it, but possible)?

If it's hosted in China, they can tell you whatever you want to hear and do whatever they want to do. What are you going to do? Take a CCP company in front of a CCP judge?

They can do the same thing in the US. What are you going to do, sue OpenAI or Anthropic?

Re: DeepSeek V4 Flash 0731

#318
post #154

The benchmark performance tells me DeepSeek v4 Flash could be very cost-effective at playing SNES/Gameboy games.

It won't be long before I can just stay home, and have my robot ride my bike for me.

I'm going to send mine to visit my mom. It's so hot in August.

Re: DeepSeek V4 Flash 0731

#319
post #266

The token price seem to be jigged, how do you know if it's subsidized or temporary. Anyone can just lower the token price to get to the left.

Almost every service provider in the AI field is subsidizing their token cost to some degree, they're all shooting for marketshare and lock-in (and they're not really achieving the latter).

Re: DeepSeek V4 Flash 0731

#320

I'm not sure about all these benchmarks, I did some very simple tests (I have my own benchmarks https://upmaru.com/llm-tests ) and these models fail, not sure if it's the inference provider or the model. They seem to be optimized for benchmarks more than real use cases. Do anything outside their distribution (even if it's not complex) they fail. I Compared Deepseek V4 Flash 0731 (low) to Gemini 3.5 Flash Lite (minima…

You're using it on low, that's why. There's a huge difference in performance from low to max effort.
Post reply on HN