Live data from Hacker News

DeepSeek V4 Flash 0731

arcprize.org

381–390 of 477 posts

Re: DeepSeek V4 Flash 0731

#381

Earlier quoted context omitted.

I think you're overlooking the fact that for long-horizon tasks, even small errors compound over time and can lead to catastrophic outcomes. For simple queries, we have reached the threshold since the beginning of the year, and models are good enough from every provider to make a meaningful difference between one another. (ChatGPT, Claude, Gemini, Grok, MuseSpark, Kimi, DeepSeek, GLM...) The real unlock will be, and…

That doesn’t make sense. It’s not like SOTA models are error free, yet we still use them. You use Fable 5 right? If that’s good enough for you now, why wouldn’t a Chinese model that’s as good as Fable 5 but at 10% the cost be good enough in 6 months?

> You use Fable 5 right? If that’s good enough for you now, why wouldn’t a Chinese model that’s as good as Fable 5 but at 10% the cost be good enough in 6 months?

This is why I quite like Kimi K3 - close to the same performance (definitely like Opus, approaching Fable), noticeably cheaper, generally good enough for me to daily drive. Only problem is that their official provider (on the Vivace plan) feels kinda slow, I'd say close to 2x slower than Opus on Max reasoning on average (probably more relatable than Fable).

Re: DeepSeek V4 Flash 0731

#382

It's really amazing to see how the gaps between the self hostable models and the closed models has been shrinking in the last 24 months. And how this has been accelerating!! I felt this very hard when I had to travel in the middle of nowhere in south america, with no network, and wanted to keep an LLM model on my macbook pro with 48GB of RAM. That was back in April 2026, a few months ago. I downloaded Google Gemma 4…

I just don't know... This post sounds like an Ai bot.

Re: DeepSeek V4 Flash 0731

#383

Earlier quoted context omitted.

If that’s the case businesses would be seeing millions to billions of profit gain (or cost reduction) in the past 4 months as they went from Opus 4.6 to Fable 5. But that’s simply not the case. It’s very clear that vast majority of the business do not generate additional value from incremental intelligence gain from these models. There is a reason why Chinese open weight models are now popular even in American enterp…

I work for a FAANG, and have my own personal projects for which I use the Chinese models. and the big models do indeed save/make us a lot of money. The Chinese models are not good enough for anything other than pair programming, which is just a very last-gen way of using agents. And when the big US models get better we will move with them. Until we stop seeing returns there is no "good enough", I don't know why this…

It depends on the use case. And most companies (like 90%+) do not have the coffers FAANG has and price does make a big difference.

Re: DeepSeek V4 Flash 0731

#384

Earlier quoted context omitted.

I think you're overlooking the fact that for long-horizon tasks, even small errors compound over time and can lead to catastrophic outcomes. For simple queries, we have reached the threshold since the beginning of the year, and models are good enough from every provider to make a meaningful difference between one another. (ChatGPT, Claude, Gemini, Grok, MuseSpark, Kimi, DeepSeek, GLM...) The real unlock will be, and…

> these compounding errors that accumulate in dumber models While SOTAs handle these errors better, they compound in all models and there's a term for that. It starts with cluster and ends with an expletive. I wish I could, but I don't see the need for human steering going away soon if the task involves anything novel (see Terry Tao's chat).

[dead]

Re: DeepSeek V4 Flash 0731

#385

Earlier quoted context omitted.

Except Fable won’t be costing $100 for enterprises that will be considering the Chinese models. If $100 Claud Max subscription works for you, then great. But you have to remember your pricing is subsidized by enterprises that pay hundreds of thousands of dollars each month, if not more, to Anthropic. For those companies, a Chinese model that can cut their AI spend from $1M/month to $200k suddenly seems attractive. An…

I don’t think the industry knows how to price this stuff. Deepseek is great (I’m running it on a RTX 6000 pro setup) but it’s nothing like Fable. It’s still strongly human-in-the-loop which is fine, until you experience how good these models can be. Think about it this way. Let’s say you could buy an LLM that gets things right 98% of the time. But there’s another LLM that’s 100x the price but gets things right 99.9%…

> but it’s nothing like Fable. It’s still strongly human-in-the-loop which is fine,

It's funny to see that Anthopic shills have been saying the exact same thing for the past two years now (and it was OpenAI fans before). It's amazing to see that Claude 3 Sonnet was "great" but now that even Qwen 9B is better than this version of Sonnet DeepSeek V4 is still not good enough despite being stronger than Opus 4.7 was.

> Let’s say you could buy an LLM that gets things right 98% of the time. But there’s another LLM that’s 100x the price but gets things right 99.9% of the time

If you think Fable makes 20 times fewer mistakes than DS4 you're delusional. It doesn't even do 20 fewer mistake than Gemma 4…

Re: DeepSeek V4 Flash 0731

#386

Earlier quoted context omitted.

To put actual numbers on it, since using AI to start solving all kinds of bottlenecks/inefficiencies in our small business, we've seen monthly net profit go up by around $4,000 USD. These are semi-permanent fixes, and the tech is only partially deployed. I am the only one using it, and I only use it part time. We've just spun up our first Hermes agent, with direct API access to our main inventory system and that's ex…

"оur first Hermes agent, with direct API access to our main inventory system" – let me assure you that absolutely nothing can go wrong here, mate. /s

Don't you think that's a bit of an inane comment to make about a setup you know nothing about?

Surely they are using read-only access.

Re: DeepSeek V4 Flash 0731

#387

Earlier quoted context omitted.

If that’s the case businesses would be seeing millions to billions of profit gain (or cost reduction) in the past 4 months as they went from Opus 4.6 to Fable 5. But that’s simply not the case. It’s very clear that vast majority of the business do not generate additional value from incremental intelligence gain from these models. There is a reason why Chinese open weight models are now popular even in American enterp…

I work for a FAANG, and have my own personal projects for which I use the Chinese models. and the big models do indeed save/make us a lot of money. The Chinese models are not good enough for anything other than pair programming, which is just a very last-gen way of using agents. And when the big US models get better we will move with them. Until we stop seeing returns there is no "good enough", I don't know why this…

> The Chinese models are not good enough for anything other than pair programming, which is just a very last-gen way of using agents.

This matches my experience with DeepSeek V4 Pro at Max reasoning, the preview version of the model kept regularly messing things up. About 30-60% of additional time to fix the output was needed.

On similar tasks, GLM 5.2 at Max reasoning screwed up maybe 20-30% of the time, while it still definitely made noticeable mistakes, they were far fewer in total and less egregious.

Kimi K3 at Max reasoning drops that value to below 10%, it's about as good as Opus or approaches Fable in some tasks. At High reasoning it also seems to be pretty close to Opus 4.8, not sure about the latest Opus model yet, but it's up there.

Only problem is that K3 is nowhere near as cheap as DeepSeek models, despite me personally liking the writing tone more (less Anthropic slop) and finding that it doesn't block my cybersecurity prompts, recently reproduced SQLi with a proof of context so I could justify fixing it.

My overall thoughts (released over some time):

https://blog.kronis.dev/blog/ai-slop-is-a-self-inflicted-tra...

https://blog.kronis.dev/blog/kimi-k3-is-out-is-anthropic-don...

https://blog.kronis.dev/blog/z-ai-s-glm-5-2-is-a-great-model...

I'd say as Chinese models get better, whatever moat Anthropic and OpenAI have dissipates. Currently the main things keeping me with Anthropic are their performance (tokens/second) and the fact that their visualization abilities within the app are pretty good.

Re: DeepSeek V4 Flash 0731

#388

Earlier quoted context omitted.

> My initial thought was to sign up for ChatGPT, but I had $20 in OpenRouter so I've been trying out DeepSeek V4 Pro with Pi for the last few days and I gotta say, it's good enough for my use case If you prefer subscriptions, OpenCode Go ($10/mo), Cline Pass ($10/mo), Atlas Code ($20/mo), and CommandCode ($1/mo) serve some of the best open weights with generous limits. OpenCode Go currently offers $120 for $10 on Dee…

> OpenCode Go currently offers $120 for $10 on DeepSeek Flash v4 At DeepSeek's absurdly low rates or market rates?

OpenCode (used to [0]?) buys inference from DeepSeek for DeepSeek models, and their $ rates reflect that: https://opencode.ai/docs/go/#usage-limits

Careful with using OpenCode's accounting for DeepSeek v4 Pro, though: https://github.com/anomalyco/opencode/issues/39822

[0] https://github.com/anomalyco/opencode/issues/39857

Re: DeepSeek V4 Flash 0731

#389

Earlier quoted context omitted.

Didn’t deepseek recently announce prices will go up significantly? Right now the US dominates everyone else in actual chips in data centers. So even if deepseek etc tries to undercut, they’re very capacity limited.

It's far from significant, it's partially doubled during peak hours. They could 16x it and it would still be two orders of magnitude better value than OAI's $200/mo plan. It's that good. They are far from capacity limited, and even if they were, you can rent a single MI300X from somewhere like Hot Aisle and get more tk/s than you'll be able to use.

How does it compare to 5.6 Luna after the permanent 80% price cut?

That one is dirt cheap at API pricing, I can't imagine quota is going to be a concern on the $200 subscription, which in my opinion easily supports full time use of 5.6 Sol on xhigh.

Re: DeepSeek V4 Flash 0731

#390
post #352
post #286

Earlier quoted context omitted.

There are more than two options. What about great engineers with great models?

Again, I was answering to: > You don't win the stock market or make the deadliest drone by switching to the cheap model The question is not "can you win with the best model?", it is "can you not win without the best model?".

[dead]
Post reply on HN