Live data from Hacker News

DeepSeek V4 Flash 0731

arcprize.org

281–290 of 474 posts

Re: DeepSeek V4 Flash 0731

#281

I've been using it extensively since the release and the best summary I can give is that it's good enough to use it for (almost) everything and cheap enough that the cost are irrelevant. I'm running it in Oh My Pi with a second instance running as "advisor" and even with 5-6 active sessions (effectively 12 streams) I'm struggling to spend more than 5 bucks per day. OpenCode Go even has double limits temporarily so fo…

And software keeps getting worst. The analogy I like is that building software is running a Michelin restaurant. The moment you scale, the chef is just writing cooking books and is absent, and you move into franchising, you will be amazed at the bottom line revenue scaling, while customers will be progressively appalled with the food...

Not that I disagree, but the average software before AI was more like a McDonald's. I have genuinely seen companies who were writing code a lot worse than what AIs produce nowadays. Doesn't necessarily mean that their software is better now, but my point is that before AI, I don't think that software compared to Michelin chefs.

Re: DeepSeek V4 Flash 0731

#282

I've been using it extensively since the release and the best summary I can give is that it's good enough to use it for (almost) everything and cheap enough that the cost are irrelevant. I'm running it in Oh My Pi with a second instance running as "advisor" and even with 5-6 active sessions (effectively 12 streams) I'm struggling to spend more than 5 bucks per day. OpenCode Go even has double limits temporarily so fo…

it is funny when people say i am struggling to spend money.

Re: DeepSeek V4 Flash 0731

#284

Earlier quoted context omitted.

That doesn’t make sense. It’s not like SOTA models are error free, yet we still use them. You use Fable 5 right? If that’s good enough for you now, why wouldn’t a Chinese model that’s as good as Fable 5 but at 10% the cost be good enough in 6 months?

I think we put up with Fable's occasional hiccups because there's nothing better at the moment. I use Claude Code semi-heavily for my small business, and the $100/mo I pay for that is a rounding error compared to the value it provides. If I can avoid spending an hour or two "massaging" the output from a lower-end model once, or it avoids introducing one load-bearing (sorry, couldn't resist) bug, then that's the entir…

Except Fable won’t be costing $100 for enterprises that will be considering the Chinese models.

If $100 Claud Max subscription works for you, then great.

But you have to remember your pricing is subsidized by enterprises that pay hundreds of thousands of dollars each month, if not more, to Anthropic.

For those companies, a Chinese model that can cut their AI spend from $1M/month to $200k suddenly seems attractive.

And unfortunately for the American tech industry, the valuation is based off those enterprise deals, not your $100/month Claude Max subscription.

Re: DeepSeek V4 Flash 0731

#285
post #162

I've been using it extensively since the release and the best summary I can give is that it's good enough to use it for (almost) everything and cheap enough that the cost are irrelevant. I'm running it in Oh My Pi with a second instance running as "advisor" and even with 5-6 active sessions (effectively 12 streams) I'm struggling to spend more than 5 bucks per day. OpenCode Go even has double limits temporarily so fo…

If what you're saying is true and accurate, then US-based AI labs are in big trouble. The only saving grace might be some sort of a 'national security' proclamation banning the use of state-of-the-art Chinese (and non-US) models across US federal and state governments and large enterprises (especially ones with federal government contracts), but even still, US AI labs will probably lose out massively on international…

> If what you're saying is true and accurate, then US-based AI labs are in big trouble.

I've been working with DeepSeek V4 Flash 0731. I'd say that it's maybe not quite as smart as Opus 4.5, but it's willing to think things through carefully and keep going until it gets a good answer. So it's a decent Opus 4.5 replacement. Just let it cook.

It isn't Opus 5 or Fable 5. But it's nearly free on Open Router, and it's self hostable on a Mac Studio with plenty of RAM, or using an RTX Pro 6000 Blackwell or two. Which is chump change for any company that employs programmers.

It would absolutely have been a frontier model last December.

Re: DeepSeek V4 Flash 0731

#286
post #279
post #268

Earlier quoted context omitted.

The bet isn't that people will be able to automatically reply on bugs and rack up API charges. The bet is on using AI to gain competitive advantage. You don't win the stock market or make the deadliest drone by switching to the cheap model

> You don't win the stock market or make the deadliest drone by switching to the cheap model Really? How many times a small team has outperformed a much bigger one just because they were "doing it right"? I have been in software companies where most software produced was bad. Not just the code, the overall design everywhere. So... bad engineers with the most expensive model, or great engineers with cheaper models?

There are more than two options. What about great engineers with great models?

Re: DeepSeek V4 Flash 0731

#287

Earlier quoted context omitted.

Seconded. I love OpenCode and Pi, but omp is my daily driver.

What's good about it? I use OpenCode and it does what I need, basically.

Spawning subagents via orchestrate and using /advisor are both super valuable. Haven't really unlocked the full capability of omp yet, but hoping to over time.

Re: DeepSeek V4 Flash 0731

#288

Earlier quoted context omitted.

I've posted a few times about my project that's a collection of 30k-250k webapps that are served from a WebDAV server. The apps know how to write updated copies of themselves back to the server. My family uses it. I have gallery apps (yearbooks for each year are a lot of fun!) of us on trips and just living, an outlining app that's a mesh of Workflowy and Org Mode (it's called Fluxtral), a markdown-backed app (it use…

Sounds fascinating! A blog write-up about your platform would be a fun read, if you're up to it

For sure! I'll be doing a Show HN at some point, just want to feel a bit more confident about certain aspects first.

Re: DeepSeek V4 Flash 0731

#289

Earlier quoted context omitted.

Sounds fascinating! A blog write-up about your platform would be a fun read, if you're up to it

For sure! I'll be doing a Show HN at some point, just want to feel a bit more confident about certain aspects first.

Makes sense, I'll look out for it! although of course most Show HNs these days get lost in a deluge of submissions.. you might actually be better off omitting the Show HN when you submit it..

Re: DeepSeek V4 Flash 0731

#290
post #276

Earlier quoted context omitted.

I think we put up with Fable's occasional hiccups because there's nothing better at the moment. I use Claude Code semi-heavily for my small business, and the $100/mo I pay for that is a rounding error compared to the value it provides. If I can avoid spending an hour or two "massaging" the output from a lower-end model once, or it avoids introducing one load-bearing (sorry, couldn't resist) bug, then that's the entir…

> I think we put up with Fable's occasional hiccups because there's nothing better at the moment. Which was an argument for using every less powerful model since the moment they got useful, right? When was that? Opus 4.5 maybe? Let's say Opus 4.5 for the sake of the argument. So back then we were like "DeepSeek is not good enough, I need Opus 4.5". Now DeepSeek is better than Opus 4.5. So if Opus 4.5 was good enough…

To put actual numbers on it, since using AI to start solving all kinds of bottlenecks/inefficiencies in our small business, we've seen monthly net profit go up by around $4,000 USD. These are semi-permanent fixes, and the tech is only partially deployed. I am the only one using it, and I only use it part time.

We've just spun up our first Hermes agent, with direct API access to our main inventory system and that's expected to find another few grand per month in misallocation/inefficiency.

I wouldn't be surprised if we were doing more like $10k/mo higher in 6-9 months' time.

When you're talking about numbers like this, the fact that one AI is $100/mo and another is $10/mo or $40/mo doesn't matter. They could make GLM-5.2, or any other Opus 4.5-class model free and it still wouldn't make sense to deploy in a commercial context.

The other angle I'd approach things from is that Opus 4.5 (and I'd agree with you that that model was the saddle point) was "good enough" for the types of things we were asking it to do back then, but as the models have become more capable the tasks we're asking them to do have also expanded with it.

I know I've personally gone from "hey can fix this race condition with a Redis mutex" 6 months ago to "Independently redesign this full embedded USB stack and QA it end-to-end, working around a specific Kernel bug in macOS Tahoe that requires decompilation to find the source of, while keeping in mind the constraints of our 8-bit AVR chip from 2011" now.

But that said, yes, maybe in 5 years' time we will reach an "intelligence saturation" where the average person won't be able to even conceive of how to use the new SOTA.

Post reply on HN