Live data from Hacker News

DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

artificialanalysis.ai

51–60 of 342 posts

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#51
post #20

Earlier quoted context omitted.

I’ve been using v4 flash for an app I’m building [1] and it’s amazing how cost effective and good it is coming from having always used gpt, opus and sonnet models. It’s so cost effective I can offer a generous free tier since my goal isn’t to make money with it. [1] https://trysojourn.app

I get where you're coming from, and the intent to make it easier for people to find examples and verses, but there's a fine line with LLMs giving you answers, is that it's interpreting it in some form. Doesn't that run counter to prevailing ideology, that you're meant to either struggle with the materials / seek understanding yourself, or have your religious leaders interpret/receive those insights?

You’re correct. The hard line for me is ensuring that any scripture presented to the user is verified correct. LLMs can’t be trusted in this regard.

One feature of the app is that all scripture is verified and what’s show to the user doesn’t come from the LLM at all and instead a trusted source.

I think exploring scripture this way does not alleviate you from struggling to learn and apply it. It hasn’t for me.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#52
post #20

Earlier quoted context omitted.

I’ve been using v4 flash for an app I’m building [1] and it’s amazing how cost effective and good it is coming from having always used gpt, opus and sonnet models. It’s so cost effective I can offer a generous free tier since my goal isn’t to make money with it. [1] https://trysojourn.app

I get where you're coming from, and the intent to make it easier for people to find examples and verses, but there's a fine line with LLMs giving you answers, is that it's interpreting it in some form. Doesn't that run counter to prevailing ideology, that you're meant to either struggle with the materials / seek understanding yourself, or have your religious leaders interpret/receive those insights?

It doesn't prevent you from going to the source and struggle with the text, nor seek expert commentary.

What's difficult and doesn't have to be with philosophy/ spirituality is to find relevant bits off situation, theme etc.

This app does that very well, LLMs are good at entity recognition.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#53

I will get downvoted, but fck it. The ban on these open models is coming within weeks, if not days. As usual, the excuse will be "national security".

Hell, the US doesn't even need to act.

I claim the CCP will wise up within 2 years, possibly much much sooner, and ban their own companies from open sourcing to prevent the Americans from acquiring the capabilities.

Despite all the nonsense claims of China distilling US models, the reality is that the Americans absolutely do distill these free Chinese models, and distillation when full logprobs are available (i.e. you have access to the weights of the model) is an order of magnitude better than when you don't.

Yes, Chinese open weight models in the short term harm US closed source model providers bottom line. In the slightly longer term, "showing your hand" and publishing both the architecture innovations and the models weights will be too dangerous for the CCP to allow. This is triply true if they can release a model that beats the Americans on most benchmarks.

I've already warned investors that this is probably the closest open weight models will ever get to closed access.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#54
post #40

I will get downvoted, but fck it. The ban on these open models is coming within weeks, if not days. As usual, the excuse will be "national security".

why would you get downvoted, that's one of the most obvious next step

Because reddit unironically has better decorum around usage of their upvote/downvote system than HN does.

People on HN downvote objectively correct information because they don't like it 24/7. There's a reason the creator of Zig left and gave the computer version of a middle finger on the way out to HN!

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#55

New Deepseek models are like Christmas for me. Really big fan of low cost API models, noone does it better than DS. Until VRAM price is low enough to run models locally, this is the way to go. The subsidized subscription model won't last, API pricing "feels" closer to a true sustainable business model.

I would bet that Deepseek API pricing is still more cost effective per token than the subscriptions. With the increase in quality Deepseek Flash just got (in my personal testing so far, it seems to have improved a lot at following instructions, and has become more proactive), there really isn’t anything that can match it in terms of cost effectiveness.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#56

If deepseek v4 flash is beating DeepSeek V4 Pro, can we expect new V4 Pro which is on par with Opus 5 in couple weeks (even better if it beats Opus)?

We can expect a new v4 pro, this was a footnote in the v4 flash announcement earlier.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#57
post #30

Earlier quoted context omitted.

Oh? What are the American model censorship tells?

Genocide in Gaza...

I find that ChatGPT isn't censoring, but it is being pretty weaselly. If you ask it "is there genocide in gaza". It will say no but also say that a lot of organizations classify it as such. It will then say "it's highly disputed".

If you poke it just a few times, however, you get to the point where it will eventually say (paraphrasing) that basically only Israel, the US state department, and the ICJ say it's not a genocide.

That is to say that it's framing it as some sort of tricky complex question when it's not. And when interrogated, it basically admits that the only people who dispute it are Israel and it's supporters.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#58

New Deepseek models are like Christmas for me. Really big fan of low cost API models, noone does it better than DS. Until VRAM price is low enough to run models locally, this is the way to go. The subsidized subscription model won't last, API pricing "feels" closer to a true sustainable business model.

[flagged]

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#59
post #40

Earlier quoted context omitted.

why would you get downvoted, that's one of the most obvious next step

Because reddit unironically has better decorum around usage of their upvote/downvote system than HN does. People on HN downvote objectively correct information because they don't like it 24/7. There's a reason the creator of Zig left and gave the computer version of a middle finger on the way out to HN!

fwiw pg said early on that downvoting for disagreement is perfectly fine: https://news.ycombinator.com/item?id=117171

commenting about voting is also something the HN guidelines warns against:

> Please don't comment about the voting on comments. It never does any good, and it makes boring reading.

https://news.ycombinator.com/newsguidelines.html

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#60

It’s exciting that a model scoring this high is dirt cheap. It’s also so inefficient, when they release the full performance numbers it’s not going to be good. One example, it takes about 3.6x more tokens to finish the same work as Gemini Flash 3.6.

Using tokens to evaluate models is an outdated approach. Cost per task is what matters. Not all tokens are created equal
Post reply on HN