Live data from Hacker News

DeepSeek V4 Flash 0731

arcprize.org

411–420 of 478 posts

Re: DeepSeek V4 Flash 0731

#411

It's really amazing to see how the gaps between the self hostable models and the closed models has been shrinking in the last 24 months. And how this has been accelerating!! I felt this very hard when I had to travel in the middle of nowhere in south america, with no network, and wanted to keep an LLM model on my macbook pro with 48GB of RAM. That was back in April 2026, a few months ago. I downloaded Google Gemma 4…

I just don't know... This post sounds like an Ai bot.

Not at all.

Re: DeepSeek V4 Flash 0731

#412

I've been using it extensively since the release and the best summary I can give is that it's good enough to use it for (almost) everything and cheap enough that the cost are irrelevant. I'm running it in Oh My Pi with a second instance running as "advisor" and even with 5-6 active sessions (effectively 12 streams) I'm struggling to spend more than 5 bucks per day. OpenCode Go even has double limits temporarily so fo…

I kept running into it looping two nights ago or else getting trapped in a reasoning loop it couldn’t escape from. Switching to Pro helped, but I ultimately had to use GLM-5.2 to recover my session. (GPT-5.6-Sol’s cybersecurity guardrails went off since the problem I was trying to fix involved a race condition where it would segfault and the words “stack frame” in my session made it decide I was being naughty. Yet an…

What was your prompt? If something benign, then it sounds like your harness has an issue with its tool calling that the agent is atrying to work around.

Re: DeepSeek V4 Flash 0731

#413

Earlier quoted context omitted.

I don’t think the industry knows how to price this stuff. Deepseek is great (I’m running it on a RTX 6000 pro setup) but it’s nothing like Fable. It’s still strongly human-in-the-loop which is fine, until you experience how good these models can be. Think about it this way. Let’s say you could buy an LLM that gets things right 98% of the time. But there’s another LLM that’s 100x the price but gets things right 99.9%…

If that’s the case businesses would be seeing millions to billions of profit gain (or cost reduction) in the past 4 months as they went from Opus 4.6 to Fable 5. But that’s simply not the case. It’s very clear that vast majority of the business do not generate additional value from incremental intelligence gain from these models. There is a reason why Chinese open weight models are now popular even in American enterp…

> It’s very clear that vast majority of the business do not generate additional value from incremental intelligence gain from these models.

It's been 2 months since Fable was released to the general, man. Nobody knows what's going on inside of these companies except the people at the coal face.

Re: DeepSeek V4 Flash 0731

#414
post #258

Earlier quoted context omitted.

Nice. Does it use a summarization, or a hard cutoff?

llama.cpp uses a hard cutoff. The agent then does "something" that is specific to the agent's implementation and configuration. It might summarize and then "finish the thought" with a different model, and then resubmit the prompt to the llama.cpp API endpoint with .. prefilled. The primary model then infers the remainder of the reply.

llama.cpp does a hard cut off on budget; it can set a reasoning-message as default but the client _can_ set a per message reasoning-message, so it's possible a smart harness could inspect the cut of thoughts and trim and do whatever.

Re: DeepSeek V4 Flash 0731

#415

Earlier quoted context omitted.

If that’s the case businesses would be seeing millions to billions of profit gain (or cost reduction) in the past 4 months as they went from Opus 4.6 to Fable 5. But that’s simply not the case. It’s very clear that vast majority of the business do not generate additional value from incremental intelligence gain from these models. There is a reason why Chinese open weight models are now popular even in American enterp…

I work for a FAANG, and have my own personal projects for which I use the Chinese models. and the big models do indeed save/make us a lot of money. The Chinese models are not good enough for anything other than pair programming, which is just a very last-gen way of using agents. And when the big US models get better we will move with them. Until we stop seeing returns there is no "good enough", I don't know why this…

> just a very last-gen way of using agents.

Fable has only been out for a month but somehow everyone is supposed to have moved to a completely different way of working that supposedly only works for Fable and nothing else…

Re: DeepSeek V4 Flash 0731

#416

Earlier quoted context omitted.

I work for a FAANG, and have my own personal projects for which I use the Chinese models. and the big models do indeed save/make us a lot of money. The Chinese models are not good enough for anything other than pair programming, which is just a very last-gen way of using agents. And when the big US models get better we will move with them. Until we stop seeing returns there is no "good enough", I don't know why this…

It depends on the use case. And most companies (like 90%+) do not have the coffers FAANG has and price does make a big difference.

Unless you are very cash strapped, Fable is a very nominal fixed cost compared to the benefit of what it offers (fully autonomous agents, and no longer needing to pair program with one).

And even it isn't "enough". I can very clearly see myself using more advanced agents to move up the abstraction ladder.

For businesses that have actual problems to solve, I see them investing in the frontier for a good bit longer, probably until we have AGI that can replace employees, maybe even a bit after.

This is why I find the "good enough" arguments silly. Like, the usefulness of an AI tops out to you when you can pair program with it? Seriously? You cannot envision ways in which more advanced AI enables you to do more, better? That's bizarre to me. I don't ever see myself running out of problems to solve.

Re: DeepSeek V4 Flash 0731

#417

Earlier quoted context omitted.

I work for a FAANG, and have my own personal projects for which I use the Chinese models. and the big models do indeed save/make us a lot of money. The Chinese models are not good enough for anything other than pair programming, which is just a very last-gen way of using agents. And when the big US models get better we will move with them. Until we stop seeing returns there is no "good enough", I don't know why this…

> just a very last-gen way of using agents. Fable has only been out for a month but somehow everyone is supposed to have moved to a completely different way of working that supposedly only works for Fable and nothing else…

This stuff is just obvious to anyone working with the latest models. Fully autonomous agents are a game changer. Having to pair program with one is indeed "last gen", I haven't done that for a month and I won't ever be doing that again in my life, outside of personal projects.

Re: DeepSeek V4 Flash 0731

#418
post #162

I've been using it extensively since the release and the best summary I can give is that it's good enough to use it for (almost) everything and cheap enough that the cost are irrelevant. I'm running it in Oh My Pi with a second instance running as "advisor" and even with 5-6 active sessions (effectively 12 streams) I'm struggling to spend more than 5 bucks per day. OpenCode Go even has double limits temporarily so fo…

If what you're saying is true and accurate, then US-based AI labs are in big trouble. The only saving grace might be some sort of a 'national security' proclamation banning the use of state-of-the-art Chinese (and non-US) models across US federal and state governments and large enterprises (especially ones with federal government contracts), but even still, US AI labs will probably lose out massively on international…

These companies know their value was never in the models. There’s a scramble to acquire as much hardware as possible, and to build as many services that people get soft locked into, such that they have a moat when the open weight models fully catch up. Doesn’t matter that the models are open weight, if you want access to the hardware you will have to pay.

I dont see how the outlook is any better for the open weight companies. They’re in the exact same situation as the closed weight companies except they have had much less revenue, and built up less of a brand, leading up the the point where they are equal in terms of model quality.

Re: DeepSeek V4 Flash 0731

#419

Earlier quoted context omitted.

Real question: is there anybody that is both maintaining alpha-dev capability by keeping abreast of all these daily changes, while also reserving enough time to actually work? Seems like we've reached the event horizon of whether AI advances are worth paying attention to.

I don't think you need to be keeping abreast of them really, you just need to be using the best model you can get enough tokens from, which for many people is Fable 5 @ $200ish, ideally fanning out implementation to cheaper models

> which for many people is Fable 5

Not for me, Fable refuses to debug Linux kernel bugs. Unless you say who you're speaking for, it sounds like you're just shilling for Anthropic.

Re: DeepSeek V4 Flash 0731

#420

Earlier quoted context omitted.

I work for a FAANG, and have my own personal projects for which I use the Chinese models. and the big models do indeed save/make us a lot of money. The Chinese models are not good enough for anything other than pair programming, which is just a very last-gen way of using agents. And when the big US models get better we will move with them. Until we stop seeing returns there is no "good enough", I don't know why this…

I use Chinese models, even smaller local ones, for much more than pair programming. If we are talking about deepseek v4 flash, which is basically a frontier model, it is much more capable than the local models I run on my MacBook Pro. The only issue really is finding the right harness. I do have a way of correcting through redundancy, though. If you are just vibe coding, you need to use the most capable model you can…

> If you are just vibe coding, you need to use the most capable model you can find and even then it might not be good enough

I mean, this proves my point. Better models enable you to get more done. With Fable, 80% of the time, I no longer have chat with an agent over the details of a PR. I give it an outcome and it gets done. This means I can work on much more with the limited time I have.

And I don't see this ending. When better models come out that take that from 80% to 99.x%, I will have that better model manage teams of other models and move up the abstraction layer.

If models get even better than that, perhaps I stop reviewing PRs entirely. Maybe normies can start using agents to build real things.

Unless your business doesn't have many problems to solve and isn't in a competitive environment, it will benefit from using the best models.

Post reply on HN