Live data from Hacker News

DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

artificialanalysis.ai

41–50 of 343 posts

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#42
It’s exciting that a model scoring this high is dirt cheap.

It’s also so inefficient, when they release the full performance numbers it’s not going to be good.

One example, it takes about 3.6x more tokens to finish the same work as Gemini Flash 3.6.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#43
post #3

404. I believe this is the correct URL: https://artificialanalysis.ai/models/deepseek-v4-flash

Yes, sorry, I went into anti-procrastination mode after I posted. I hope someone fixes it.

@dang is that how we call you?

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#44
post #30
post #22

Earlier quoted context omitted.

Western models censor just as much shit as the Chinese models do, big guy, it’s just different material. While we should be pushing for universal fully uncensored models, this comment is lazy and trite at this point. But you already know that.

Oh? What are the American model censorship tells?

Genocide in Gaza...

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#45

what a horribly heavy and resource-consuming website...

we need a benchmark website benchmark

A benchmark website to benchmark benchmark websites?

Or a benchmark to benchmark benchmarks?

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#48

Earlier quoted context omitted.

Yes, sorry, I went into anti-procrastination mode after I posted. I hope someone fixes it.

@dang is that how we call you?

no, like this: hn@ycombinator.com

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#50
Daily reminder that none of these numbers are valid in a world where no one publishes the sampling settings used.

Daily reminder that improving your samplers from the garbage default top_p/top_k to min_p or subsequent methods dramatically improves the performance of these models, and makes most quantities like measured "verbosity" and subsequent calculations of "intelligence per token" meaningless

Daily reminder that no one, including within academic AI research, AI engineers, normies, etc takes LLM sampling seriously enough.

Post reply on HN