Live data from Hacker News

DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

artificialanalysis.ai

231–240 of 342 posts

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#231

Earlier quoted context omitted.

The problem is picking between models. I do not want to spend my time switching models and trying to decipher which model should be used for what. Maybe that's just a me problem that I need to figure out.

DeepSeek v4 is honestly good enough that I'm fine throwing it at everything in my hobby projects. I guess now I'll be switching from V4 Pro-Preview to V4 Flash. My only real complaint is that they can't do images, which limits their ability to autonomously debug some kinds of issues Of course you can get more bang for your buck by being more deliberate. But that's equally true with US frontier models. You can optimiz…

CodeWhale is a coding agent that auto-routes requests to Flash / Pro based on complexity, as determined by Flash. It's also tuned for DeepSeek's caching behavior, making things even more inexpensive. I'm retired, but I've been using it for just over two months at about the rate I would use it if I were working half-time, and I've spent $19 total.

https://github.com/Hmbown/CodeWhale

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#232

Somewhat relatedly, how do the economics for Huggingface work? They must be hosting petabytes of models and datasets by now. I have downloaded quite a few “just in case”, only to replace them with the later iteration months later. Does the file hosting actually cost peanuts when you do it yourself and the cloud has shattered my understanding of what it actually costs to deliver so much data?

File hosting is pretty cheap. Egress traffic is about $90/TB in the cloud, but around $1/TB in the real world. Storage is in the realm of $5/TB/month after adding redundancy At the scale of Huggingface, that still amounts to a lot load of money. Significantly less than if you did the same in AWS, but still a lot That said, they do have a deal with AWS to make the data available in AWS ip space. Maybe they got some ch…

It's less than $1/TB. If you have settlement-free peering it's like fractions of a penny at scale. Cloud provider DTO pricing is literally the biggest scam on Earth.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#235

Earlier quoted context omitted.

Hell, the US doesn't even need to act. I claim the CCP will wise up within 2 years, possibly much much sooner, and ban their own companies from open sourcing to prevent the Americans from acquiring the capabilities. Despite all the nonsense claims of China distilling US models, the reality is that the Americans absolutely do distill these free Chinese models, and distillation when full logprobs are available (i.e. yo…

You’re describing the dump and pump strategy that china has historically used across a number of industries. Stands to reason that this is what is going on.

[deleted]

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#236
post #20

Earlier quoted context omitted.

I’ve been using v4 flash for an app I’m building [1] and it’s amazing how cost effective and good it is coming from having always used gpt, opus and sonnet models. It’s so cost effective I can offer a generous free tier since my goal isn’t to make money with it. [1] https://trysojourn.app

I get where you're coming from, and the intent to make it easier for people to find examples and verses, but there's a fine line with LLMs giving you answers, is that it's interpreting it in some form. Doesn't that run counter to prevailing ideology, that you're meant to either struggle with the materials / seek understanding yourself, or have your religious leaders interpret/receive those insights?

> Doesn't that run counter to prevailing ideology, that you're meant to either struggle with the materials / seek understanding yourself, or have your religious leaders interpret/receive those insights?

I think it depends, Catholics wouldn't be able to use this because the Magisterium is the ultimate authority on interpreting Scripture, so the personal interpretation isn't really needed. This is not to say that Catholics don't read the Bible, they are encouraged to do so since it deepens their faith

On the other hand, for Protestant it varies, the High Church denominations are closer to Catholics (though none of them accept the Magisterium) in terms of scripture interpretation, but the Low Church ones (like Baptists or Non-Denominational ) are more open to personal interpenetration.

Disclaimer: I'm a Catholic, so if I made a mistake here fellow Protestants, please correct me.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#237

Earlier quoted context omitted.

The leather bag designs are shipped from Paris or Milan to China for production.

But some worker in Italy bolts on a zipper in the end so it’s actually Italian, believe it or not.

That worker is also Chinese, but in Milano.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#238

It’s exciting that a model scoring this high is dirt cheap. It’s also so inefficient, when they release the full performance numbers it’s not going to be good. One example, it takes about 3.6x more tokens to finish the same work as Gemini Flash 3.6.

> inefficient

That depends. Is it also more reliable?

If two books, one big one slim, prove the same thesis, what I would be interested in is the quality of the content, not the size. There can be a measure of efficiency in "have you really thought it through", but it is clearly complex - it requires measuring how solid the reasoning is.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#239

Earlier quoted context omitted.

They announced it already, read the tech report. "For the Code Agent tasks among the public benchmarks above, DeepSeek-V4-Flash-0731 is evaluated with the minimal mode of DeepSeek Harness (to be released) as the agent framework , using the max reasoning effort level with temperature = 1.0, top_p = 0.95."

> They announced it already I meant something I could download and run.

Isn't Reasonix [0] DeepSeek's own harness? Or, are they building a new one?

[0] https://reasonix.io/

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#240
post #221

Earlier quoted context omitted.

My conspiracy theory is that this is the new space race, and the CCP encourages this to show the world what Chinese engineers are capable of, and tank the Anthropic/OpenAI valuation bubble as a desirable side effect.

Oh no, 1kkk market country with top-tier research labs developing its own technology, must be evil!

Well, no, they're not aligned with our interests therefore we're concerned with their mastery here. You have the clause the wrong way around.
Post reply on HN