Live data from Hacker News

DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

artificialanalysis.ai

291–300 of 342 posts

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#291

Wow this thing just blew out the pareto front for intelligence/$

If Luna hadn't price dropped yesterday it would have been a real blowout.

I wonder if OpenAI knew and deliberately price dropped to front-run the news cycle in their favor.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#292

Earlier quoted context omitted.

I would like to see your "home"

Two (linked) DGX Sparks would do it I guess. Though probably slowly (I'd guess 15-20 tok/sec for decode, but higher for prefill). So ~$8-9k USD at current RAM prices, substantially less if they ever (sigh) drop. Electricity use would actually be relatively modest. But it makes little to no sense as long as API prices are what they are. Except for maybe privacy reasons.

It's at least 2.5x that speed for dual sparks and prefill is good as well. Basically going on vibes it is faster seeming than what one gets by default with openAI or Anthropic.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#293
post #284

Earlier quoted context omitted.

"The thing" isn't free [0]: It is cheap by US/EU income standards. > who is the customer If no one else, then definitely those that can only budget $1/mo to $5/mo. Probably 100s of millions, if not billions, of the Global South. [0] Free access to DeepSeek v4 Flash is indeed available from providers like OpenRouter, OpenCode, and Freebuff (to name a few), but no ZDR.

I mean, if you're paying this little, you aren't the customer; the service is either subsidized by investor money (best option), the government (which makes it unfair competition at best ) or profit is being realized elsewhere (where?)

There are dozens of providers, the model is just really cheap to run thanks to its clever architecture that optimize the compute and memory usage even in long contexts.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#294

Earlier quoted context omitted.

If only they would let you opt out of training use, it might actually be a viable option.

The training use is probably responsible for this improvement tho.

They could certainly still achieve that even if they had a slightly more expensive tier with a (pinky promise) "no training" bit set, or even if it were just an opt-in most users wouldn't bother with.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#295
post #214

Now let's see Dario's price cut.

I think, it's the end of Anthropic. They can't compete, they have bills. By the end of the year, if they can't react, it's game over. Maybe US clients could be a little patriotic here, but money is money. They won't give them free money forever.

yep. I've been saying for month that exactly this will happen with DeepSeek and Kimi K3, and margins will erode, and Anthropic will fail to IPO.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#296
post #180

I have updated OpenAI's chart[1] from yesterday to include one more datapoint: DeepSeek V4 Flash 0731. It's on the frontier. https://files.parasmittal.com/openai_aa_luna_dsflash.svg 1: https://openai.com/index/advancing-the-price-performance-fro...

LOL https://artificialanalysis.ai/models/deepseek-v4-flash#intel...

High hopes for V4 Pro

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#297

I was writing a benchmark for my own harness, and DS4 flash answers as well as Fable 5 on any query. The specific agent is focused on getting precise and on point answers about a codebase. The starting point was nowhere near. E.g. asked why was X implemented in a certain way it would give bogus answers when the real answer was that there was no reason at all. The benchmark included more than 50 questions or different…

What does your harness do to squeeze that value out of DS4 flash if you don't mind sharing? It'd interesting how it compares to other harnesses (even if it's a qualitative assessment instead of quantitative).

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#298

[flagged]

You're clearly not an American and they will never catch up on hardware. I'm mostly curious why you are so excited for China to "crush" American companies?

American ultracapitalism seems to be approaching its final stages. That end can’t come quickly enough. If China helps speed it along, that’s a good thing for all of humanity.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#299

Earlier quoted context omitted.

It’s not outdated at all to use tokens to estimate performance, it’s directly related.

But why should I care? If my metrics are speed and cost? How many tokens it takes as a user is arbitrary to some extent.

If speed is a metric for you, tokens required to solve a problem affects that metric.

All else being equal passing triple the amount of tokens through a model to solve the same problem makes it slower.

Doesn’t mean this model is bad, and it has to be considered how cheap it is, but it’s a factor.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#300
post #138

Earlier quoted context omitted.

why do they do if their employers are paying for it?

because once they run out of tokens they can’t do their jobs anymore

This is a sad state of affair. I'm faster with the LLM, but I sure as hell can do everything just like I could, as before the current state of things.
Post reply on HN