Live data from Hacker News

DeepSeek V4 Flash 0731

arcprize.org

491–500 of 506 posts

Re: DeepSeek V4 Flash 0731

#491

Earlier quoted context omitted.

This would be more convincing if those providers had converged on a number that was not the exact pricing of DeepSeek themselves. Clearly DeepSeek is setting the price here and without them holding it down I expect increases.

$0.14 is CN¥1. that's where the 0.14 comes from

OK? Maybe DeepSeek set their price that way, but the majority of these OpenRouter inference providers are in the US, not China, and have no reason to care about CNY at all.

Re: DeepSeek V4 Flash 0731

#492

Earlier quoted context omitted.

If it's hosted in China, they can tell you whatever you want to hear and do whatever they want to do. What are you going to do? Take a CCP company in front of a CCP judge?

They can do the same thing in the US. What are you going to do, sue OpenAI or Anthropic?

Anthropic paid out a lot of money for the copyrighted material they used.

(not commenting on whether it was a fair amount - just saying these companies are not immune to lawsuits)

Re: DeepSeek V4 Flash 0731

#493

Earlier quoted context omitted.

It depends on the use case. And most companies (like 90%+) do not have the coffers FAANG has and price does make a big difference.

Unless you are very cash strapped, Fable is a very nominal fixed cost compared to the benefit of what it offers (fully autonomous agents, and no longer needing to pair program with one). And even it isn't "enough". I can very clearly see myself using more advanced agents to move up the abstraction ladder. For businesses that have actual problems to solve, I see them investing in the frontier for a good bit longer, pr…

> [...] probably until we have AGI that can replace employees, maybe even a bit after.

Is this what you are looking forward to?

Re: DeepSeek V4 Flash 0731

#494
post #164

Note this is the 07/31 release of DSv4 flash and not the "preview" that they put out a couple months or so ago. I've been running this model locally for a week, and the preview version before that. This updated one feels like a whole tier up. It's very capable for debugging and analyzing documents/data I upload. The killer feature, IMO, is the speed. On 2x RTX Pro 6000 Blackwell, its ~8k tok/s prefill and ~250 tok/s…

I wish I could say that it performs reasonably on my hardware. 8x Radeon AI Pro 9700XTs, and I can't get it to hit double-digit tokens per second. Neither vllm nor llama.cpp, with various combinations of quants, draft models, and parallelism methods can get it to run at a tolerable speed. I'll be sticking with StepFun 3.7-Flash for the foreseeable future :(

Re: DeepSeek V4 Flash 0731

#495

Earlier quoted context omitted.

I don't think you need to be keeping abreast of them really, you just need to be using the best model you can get enough tokens from, which for many people is Fable 5 @ $200ish, ideally fanning out implementation to cheaper models

> which for many people is Fable 5 Not for me, Fable refuses to debug Linux kernel bugs. Unless you say who you're speaking for, it sounds like you're just shilling for Anthropic.

I would love to be shilling for Anthropic, but I am not. I am part of a group of about 30 developers, and 80% of them are using Fable 5 and very sold on it, with the remainder being committed to Sol. Both are competent, but among our set (who will try anything), Fable 5 is definitely winning. The fucking refusals for security work are insane though, and I hate them.

I use Sol and Grok 4.5 as my inline debuggers/reviewers, and both do well, and are decent at token save. DeepSeek V4 Flash 0731 found some interesting bugs when I tried it a few days ago, and I'm curious to see if that also joins the code-review line up

Re: DeepSeek V4 Flash 0731

#497

Perhaps it might be interesting: a latent thinking version is here https://huggingface.co/nmitchko/DeepSeek-V4-Flash-0731-Laten... Does no thinking emissions for context saving.

This deserves its own hn post!

https://news.ycombinator.com/item?id=49230550

Re: DeepSeek V4 Flash 0731

#498
post #438

Did anyone else experience a change in verbosity? I've been playing around with an agent that holds your hand in a Jupiter notebook and it felt like it started writing essays versus nice, concise, helpful paragraphs like before. My gut was correct because I checked my Deepinfra usage and it was almost a 2x out-token usage for every in-token. Not a huge deal since it's still cents per session, but my bigger issue was…

[flagged]

Re: DeepSeek V4 Flash 0731

#499

Earlier quoted context omitted.

Real question: is there anybody that is both maintaining alpha-dev capability by keeping abreast of all these daily changes, while also reserving enough time to actually work? Seems like we've reached the event horizon of whether AI advances are worth paying attention to.

Alpha dev?

Made-up term. Like an alpha male or apex predator, the kind that are 10x and companies are often created around.

Re: DeepSeek V4 Flash 0731

#500
post #453

Earlier quoted context omitted.

Subsidy doesn't matter once the weights are released. You're conflating subsidized training and research with subsidized inference. DeepSeek's (allegedly government-) subsidized inference is coming to an end soon per their docs, but that's fine because the weights are open. You can run this on-prem with very modest hardware. And you can fine-tune it to your business with your data.

Then why are the major American labs in "deep trouble"? They can also continue to be subsidized and release the weights if they want to, if that's what you think it takes for survival. Seems like a low bar. The point is it takes money to keep developing models. Everyone is playing by the same rules. At this point, the US labs are trying to build businesses. I'm not really sure what the Chinese labs goals are. But I d…

They are in trouble because they offer something 5% better for 2000% the price.
Post reply on HN