Live data from Hacker News

DeepSeek V4 Flash 0731

arcprize.org

451–460 of 478 posts

Re: DeepSeek V4 Flash 0731

#451
post #162

I've been using it extensively since the release and the best summary I can give is that it's good enough to use it for (almost) everything and cheap enough that the cost are irrelevant. I'm running it in Oh My Pi with a second instance running as "advisor" and even with 5-6 active sessions (effectively 12 streams) I'm struggling to spend more than 5 bucks per day. OpenCode Go even has double limits temporarily so fo…

If what you're saying is true and accurate, then US-based AI labs are in big trouble. The only saving grace might be some sort of a 'national security' proclamation banning the use of state-of-the-art Chinese (and non-US) models across US federal and state governments and large enterprises (especially ones with federal government contracts), but even still, US AI labs will probably lose out massively on international…

One random says they are using DeepSeek and you proclaim it's all over for the top AI labs. Brilliant.

First of all, no one knows the "true cost" of any of this, yet, but we know it's expensive. To what extent are the Chinese labs being subsidized? Are they real businesses?

Second, the Chinese labs aren't some "super geniuses", while the American labs are full of clowns. As of today, like the past 3 years, American labs are SOTA. That might change, but let's not act like the American labs don't know what they are doing.

The idea that people are going to use cheaper models for cheaper work isn't some novel revelation, it's completely obvious. People are doing it already, eschewing Fable.

Re: DeepSeek V4 Flash 0731

#453
post #162

Earlier quoted context omitted.

If what you're saying is true and accurate, then US-based AI labs are in big trouble. The only saving grace might be some sort of a 'national security' proclamation banning the use of state-of-the-art Chinese (and non-US) models across US federal and state governments and large enterprises (especially ones with federal government contracts), but even still, US AI labs will probably lose out massively on international…

One random says they are using DeepSeek and you proclaim it's all over for the top AI labs. Brilliant. First of all, no one knows the "true cost" of any of this, yet, but we know it's expensive. To what extent are the Chinese labs being subsidized? Are they real businesses? Second, the Chinese labs aren't some "super geniuses", while the American labs are full of clowns. As of today, like the past 3 years, American l…

Subsidy doesn't matter once the weights are released. You're conflating subsidized training and research with subsidized inference. DeepSeek's (allegedly government-) subsidized inference is coming to an end soon per their docs, but that's fine because the weights are open. You can run this on-prem with very modest hardware. And you can fine-tune it to your business with your data.

Re: DeepSeek V4 Flash 0731

#454

I've been using it extensively since the release and the best summary I can give is that it's good enough to use it for (almost) everything and cheap enough that the cost are irrelevant. I'm running it in Oh My Pi with a second instance running as "advisor" and even with 5-6 active sessions (effectively 12 streams) I'm struggling to spend more than 5 bucks per day. OpenCode Go even has double limits temporarily so fo…

I try to use all the intelligence I can, which fable, opus and then second tier models.

I am not sure why you wouldn't want to use the SOTA models unless speed is a concern. Otherwise you are leaving quality on the table.

Re: DeepSeek V4 Flash 0731

#455

Earlier quoted context omitted.

I think you're overlooking the fact that for long-horizon tasks, even small errors compound over time and can lead to catastrophic outcomes. For simple queries, we have reached the threshold since the beginning of the year, and models are good enough from every provider to make a meaningful difference between one another. (ChatGPT, Claude, Gemini, Grok, MuseSpark, Kimi, DeepSeek, GLM...) The real unlock will be, and…

That doesn’t make sense. It’s not like SOTA models are error free, yet we still use them. You use Fable 5 right? If that’s good enough for you now, why wouldn’t a Chinese model that’s as good as Fable 5 but at 10% the cost be good enough in 6 months?

"isn't the old version good enough?"

You can say this about literally every product we buy. And yet...

Re: DeepSeek V4 Flash 0731

#456
Caching makes a huge difference to cost. On Fireworks AI, for example, if it hits the cache, you pay only 20%. And uncached is just $0.14/M tokens for DSV4-0731! I get entire re-architecture projects (with new tests and documentation) done for mere dollars. DSV4-0731 is a daily driver for me.

But note that you have to use Cline (or other harness) if using vscode. I was shocked at how poor the recent versions of GitHub Copilot are at using the cache (with Fireworks AI, but I believe it's a more generic problem).

https://x.com/vijucat/status/2085415745144672492?s=20

Re: DeepSeek V4 Flash 0731

#457
post #453

Earlier quoted context omitted.

One random says they are using DeepSeek and you proclaim it's all over for the top AI labs. Brilliant. First of all, no one knows the "true cost" of any of this, yet, but we know it's expensive. To what extent are the Chinese labs being subsidized? Are they real businesses? Second, the Chinese labs aren't some "super geniuses", while the American labs are full of clowns. As of today, like the past 3 years, American l…

Subsidy doesn't matter once the weights are released. You're conflating subsidized training and research with subsidized inference. DeepSeek's (allegedly government-) subsidized inference is coming to an end soon per their docs, but that's fine because the weights are open. You can run this on-prem with very modest hardware. And you can fine-tune it to your business with your data.

Then why are the major American labs in "deep trouble"? They can also continue to be subsidized and release the weights if they want to, if that's what you think it takes for survival. Seems like a low bar.

The point is it takes money to keep developing models. Everyone is playing by the same rules. At this point, the US labs are trying to build businesses. I'm not really sure what the Chinese labs goals are. But I do know they aren't doing charity work.

Re: DeepSeek V4 Flash 0731

#458
post #136
post #118

Earlier quoted context omitted.

How is $5/day irrelevant? In the $150/mo range you can get effectively unlimited usage of GPT 5.6 Sol (Pro plan). Why use a much weaker model for the same price?

> In the $150/mo range you can get effectively unlimited usage of GPT 5.6 Sol (Pro plan) Not true. Sol on XHigh or Max runs out even on the $200/mo plan. It's not close to effectively unlimited. Maybe at 2x the current allowance it can.

What are you doing where you're getting Sol on XHigh or Max to run out on the $200 plan? Last night, I had 48% of my weekly limit left, so I spun up 4 projects I had been working on and ran them all on Ultra with Fast mode and /goal, and it took over 4 hours to burn through that. I mean, I was reeeeally trying to use it up because I just can't use it up with normal usage for weekend and after-work programming. I've had Max crunching on something this morning for 3 hours now and I've only used 4% of my weekly limit ...

Re: DeepSeek V4 Flash 0731

#459
we got this running on 4 RTX Pro 6000's and for single request we're getting around 250 tok/s we can support about 48 concurrent requests we're seeing around 2400 agg tok/s peaking around 24-31 concurrent users. Model performance feels like gpt 5.4 - mostly using it with pi agent. the only thing i'm missing with this model is vision and i see some folks have done some work like https://huggingface.co/webbrain-one/DeepSeek-V4-Flash-0731-V... but have not yet tried it out.

Re: DeepSeek V4 Flash 0731

#460

Earlier quoted context omitted.

To put actual numbers on it, since using AI to start solving all kinds of bottlenecks/inefficiencies in our small business, we've seen monthly net profit go up by around $4,000 USD. These are semi-permanent fixes, and the tech is only partially deployed. I am the only one using it, and I only use it part time. We've just spun up our first Hermes agent, with direct API access to our main inventory system and that's ex…

I think we're even starting to reach that saturation point now for a lot of people. In my industry (law) plenty of people have tried CoPilot once or twice, or tried ChatGPT a year ago, and as a result have basically dismissed AI as being useless. The setup required to be able to get it to do end to end tasks to your liking is also substantially more work than most people are willing to put in.

I find 500x returns on my $20 Chat GPT and Claude subscriptions as I litigate pro se against large law firms in federal court.
Post reply on HN