Imagine if they had GPU resources of western labs.
DeepSeek V4 Flash 0731
61–70 of 481 posts
Re: DeepSeek V4 Flash 0731
#62It's always fun when Max reasoning is cheaper than High reasoning.
Tell your PjM who should tell your PgM who should tell your PdM, all the PMs...
Maybe if "the business" sees it is true of LLMs, they might believe it's true of giving better context to engineers up front then giving them time to think and prototype (thinking tokens are an answer prototype).
Re: DeepSeek V4 Flash 0731
#63Earlier quoted context omitted.
Dax (from Opencode) has tweeted that they can replicate or beat the price with rented GPUs. Deepseeks secret sauce is the incredibly cheap caching (magnitude cheaper than other providers). vLLM has recently released a similar approach. It's not as effective as what DeepSeek does but still an interesting development. I have no doubt that in due time other providers will match or perhaps even beat the current DeepSeek…
> Deepseeks secret sauce is the incredibly cheap caching (magnitude cheaper than other providers). Can anyone working at one of the main US labs (Google, OpenAI, Anthropic) comment on WTF they haven't even tried MLA - despite the obvious massive advantages? I know enough to know they aren't completely incompetent. So there must be a quite good reason. But it remains a mystery to me. DeepSeek's MLA is like almost 2 ye…
The big US labs are opaque and don't publish much of any technical details anymore. We don't know what they are or aren't doing, honestly.
Re: DeepSeek V4 Flash 0731
#64Why wasn’t this run against ARC-AGI-3? Or did it fail to solve anything?
Re: DeepSeek V4 Flash 0731
#65I've been using it extensively since the release and the best summary I can give is that it's good enough to use it for (almost) everything and cheap enough that the cost are irrelevant. I'm running it in Oh My Pi with a second instance running as "advisor" and even with 5-6 active sessions (effectively 12 streams) I'm struggling to spend more than 5 bucks per day. OpenCode Go even has double limits temporarily so fo…
I hadn't really thought about this but AI may well be the technology that disrupts and ultimately destroys social media.
The value proposition of something like FB or IG is, as we know, the network effect. The platform gets to extract value from user generated content. I believe that users should own the platform, a bit like the Wikimedia Foundation, because they're the ones that create value. Federation is a popular belief on HN and I've come to believe that's simply the wrong solution to the right problem.
Anyway, how these social media companies make money is by optimizing the feed for engagement. People know it too so you see people trying to build an audience by rage baiting. And then more time spent equals more advertising revenue.
But what happens when the AI can simply slurp all the posts and then filter and rank them? It destroys the engagement and advertising model. And I'm not opposed to that, honestly. It may be on eof the few good thing sto come out of AI.
Re: DeepSeek V4 Flash 0731
#66Earlier quoted context omitted.
DeepSeek V4 Flash 0731 is an open-weights model which means price is determined by competition/invisible hand of the marketplace: https://openrouter.ai/deepseek/deepseek-v4-flash-0731 With the exception of cache costs, all providers have similar input/output costs.
Not counting the cost of making the model, which is subsidized by… someone? The chinese gov i think?
Re: DeepSeek V4 Flash 0731
#67Earlier quoted context omitted.
Is it still cheaper than Luna if using an OpenAI subscription? My gut is no, but I have not done the math.
You'd have to compare against something like the OpenCode Go subscription, and I'm fairly sure deepseek napkins out cheaper in that scenario
Re: DeepSeek V4 Flash 0731
#68Earlier quoted context omitted.
Yeah either the benchmark isn't very useful anymore or V4 Flash is a really, really good model.
In my use, DeepSeek v4 Flash (which replaced the quite excellent MiniMax M3) lags behind GLM 5.2 & Muse Spark 1.2 (let alone Kimi K3). Also, K3 is a much bigger multi-modal model, while Flash is text-only and likely optimised for coding tasks.
Spark is actually the interesting one imo. It's significantly better, also significantly faster. If you are ok with letting Meta soak up your data (which DS does too) it's also the same price.
Re: DeepSeek V4 Flash 0731
#69I’ve been refreshing hacker news constantly for a week now waiting for v4 pro, after they stated it would follow «soon». I have learnt «soon» is a matter of definition.
Re: DeepSeek V4 Flash 0731
#70The DeepSeek team is so strong, very impressive. Imagine if they had GPU resources of western labs.
SV companies get way too comfortable when they have enough in the bank to stay running more than three months.