Live data from Hacker News

DeepSeek V4 Flash 0731

arcprize.org

371–380 of 477 posts

Re: DeepSeek V4 Flash 0731

#371
post #162

I've been using it extensively since the release and the best summary I can give is that it's good enough to use it for (almost) everything and cheap enough that the cost are irrelevant. I'm running it in Oh My Pi with a second instance running as "advisor" and even with 5-6 active sessions (effectively 12 streams) I'm struggling to spend more than 5 bucks per day. OpenCode Go even has double limits temporarily so fo…

If what you're saying is true and accurate, then US-based AI labs are in big trouble. The only saving grace might be some sort of a 'national security' proclamation banning the use of state-of-the-art Chinese (and non-US) models across US federal and state governments and large enterprises (especially ones with federal government contracts), but even still, US AI labs will probably lose out massively on international…

I've been using DeepSeek exclusively since I've been playing with LLMs, largely due to the price.

I joined a company that is an Anthropic shop and I am genuinely shocked.

Sonnet 5 is a little better on long horizon tasks and headless unsupervised agent workflows - but for in-IDE workflows, it's virtually unusable.

I am so used to flipping around my codebase at warp speed with DeepSeek flash. It's so fast and accurate I don't have time for parallel agents. It's a really rewarding workflow.

Moving to Sonnet, you ask is something simple like "split this into a seperate file" "implement this method" "this is my schema, implement a repository for it". It'll spend 30 minutes thinking and charge like $12. And no token caching, what are you even doing Anthropic?

It's unusable.

DeepSeek are in a league of their own

Re: DeepSeek V4 Flash 0731

#372
post #276

Earlier quoted context omitted.

> I think we put up with Fable's occasional hiccups because there's nothing better at the moment. Which was an argument for using every less powerful model since the moment they got useful, right? When was that? Opus 4.5 maybe? Let's say Opus 4.5 for the sake of the argument. So back then we were like "DeepSeek is not good enough, I need Opus 4.5". Now DeepSeek is better than Opus 4.5. So if Opus 4.5 was good enough…

To put actual numbers on it, since using AI to start solving all kinds of bottlenecks/inefficiencies in our small business, we've seen monthly net profit go up by around $4,000 USD. These are semi-permanent fixes, and the tech is only partially deployed. I am the only one using it, and I only use it part time. We've just spun up our first Hermes agent, with direct API access to our main inventory system and that's ex…

I think HN doesn't really understand fixed costs, I spend a few hundred dollars per day on Fable and the costs are irrelevant compared to what we make.

Re: DeepSeek V4 Flash 0731

#373

Earlier quoted context omitted.

I don’t think the industry knows how to price this stuff. Deepseek is great (I’m running it on a RTX 6000 pro setup) but it’s nothing like Fable. It’s still strongly human-in-the-loop which is fine, until you experience how good these models can be. Think about it this way. Let’s say you could buy an LLM that gets things right 98% of the time. But there’s another LLM that’s 100x the price but gets things right 99.9%…

If that’s the case businesses would be seeing millions to billions of profit gain (or cost reduction) in the past 4 months as they went from Opus 4.6 to Fable 5. But that’s simply not the case. It’s very clear that vast majority of the business do not generate additional value from incremental intelligence gain from these models. There is a reason why Chinese open weight models are now popular even in American enterp…

I work for a FAANG, and have my own personal projects for which I use the Chinese models. and the big models do indeed save/make us a lot of money.

The Chinese models are not good enough for anything other than pair programming, which is just a very last-gen way of using agents.

And when the big US models get better we will move with them. Until we stop seeing returns there is no "good enough", I don't know why this is so hard for HN to understand.

Re: DeepSeek V4 Flash 0731

#374

I've been using it extensively since the release and the best summary I can give is that it's good enough to use it for (almost) everything and cheap enough that the cost are irrelevant. I'm running it in Oh My Pi with a second instance running as "advisor" and even with 5-6 active sessions (effectively 12 streams) I'm struggling to spend more than 5 bucks per day. OpenCode Go even has double limits temporarily so fo…

[deleted]

Re: DeepSeek V4 Flash 0731

#375

Earlier quoted context omitted.

GPT 5.6 Luna is an extremely cheap and still very capable model. A chinese model being in the same ballpark of capability at half the price sounds believable to me.

It's significantly worse than Luna and quite a bit slower in some fairly involved tests I run.

That's fascinating, it's WAY better than luna ime. What sort of things are you testing it for?

Re: DeepSeek V4 Flash 0731

#376

Earlier quoted context omitted.

You're using it on low, that's why. There's a huge difference in performance from low to max effort.

I’m comparing same / similar settings between models. I can’t use high on one and low on others it’s not a fair test. Not sure why I was downvoted. But seems the downvoter is quick to downvote anything that doesn’t fit the narrative they’re looking for. I’m just reporting my findings.

Comparing at similar thinking hasn't much value. You can compare the tiers that have the closest price, that would be more interesting.

Re: DeepSeek V4 Flash 0731

#377

Earlier quoted context omitted.

You're using it on low, that's why. There's a huge difference in performance from low to max effort.

I’m comparing same / similar settings between models. I can’t use high on one and low on others it’s not a fair test. Not sure why I was downvoted. But seems the downvoter is quick to downvote anything that doesn’t fit the narrative they’re looking for. I’m just reporting my findings.

I'm not sure it's a fair test either to compare the "low" setting of one model with the "low" setting of another. They're completely different settings that just happen to have the same name.

Re: DeepSeek V4 Flash 0731

#378
post #376

Earlier quoted context omitted.

I’m comparing same / similar settings between models. I can’t use high on one and low on others it’s not a fair test. Not sure why I was downvoted. But seems the downvoter is quick to downvote anything that doesn’t fit the narrative they’re looking for. I’m just reporting my findings.

Comparing at similar thinking hasn't much value. You can compare the tiers that have the closest price, that would be more interesting.

Yes, I think a proper comprehensive test would be a better judge of the outcome. I may do round 2 given my first batch of models is already outdated.

Re: DeepSeek V4 Flash 0731

#380
post #170

Earlier quoted context omitted.

It is true. I don't care about having infinite frontier-level intelligence, and I don't care if Fable can one-shot frobnicate a klaxelzorp with a benchmark performance of 97%. I doubt most people do, in fact. I just want something that meets the baseline level of intelligence needed to be a really, really good pair programming agent. It shouldn't have any silly dealbreaker issues involving laziness or hallucinations,…

I wonder when we crossed the "99 percentile of intelligence for 99% of the usecases" threshold. At this point, the gains seem to be right at the very edge of bleeding edge for narrow and specialized use cases, and wonder if it'll be a sort of diminishing return from here on.

> narrow and specialized use cases

Such as Decision Making. /s

You just can't set a high enough threshold of intellectual effort for critical decisions.

Post reply on HN