Live data from Hacker News

DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

artificialanalysis.ai

71–80 of 342 posts

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#71
post #12

Does it already know the answer to what happen at Tiananmen Square? Or still avoiding it?

How often are you asking an LLM about this? You know you can just google it, right?

I have to admit it rarely comes up in the coding tasks I usually give to LLMs.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#72
post #22

Earlier quoted context omitted.

Western models censor just as much shit as the Chinese models do, big guy, it’s just different material. While we should be pushing for universal fully uncensored models, this comment is lazy and trite at this point. But you already know that.

This is a straightforward false equivalency. “Western” models do not censor in the same way, nor for the same reasons, that the Chinese models do. “Just as much” is not remotely plausible, yet it’s doing all the heavy lifting.

The Anthropic and OpenAI models are much more censored and in ways that directly prevent them to be useful, e.g. by refusing to reply to elementary questions of biology and chemistry.

Any normal user is much more likely to ask questions to which the Anthropic and OpenAI models do not answer, than to ask questions about the modern Chinese history, to which a Chinese LLM will not answer.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#74

Earlier quoted context omitted.

How exactly will they ban them?

By making companies using them "toxic" to touch. For example: no government contract to any company who uses even one vendor in it's entire chain of dependencies, who uses such open models. They can extend this further by laying more conditions, such as: any company dealing in this-this field can only use models "officially" approved as "safe". Rest you can guess how easy it would be to get that "safe" rating for suc…

Can't wait to distill Deepseek v4 flash to America-1

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#77

New Deepseek models are like Christmas for me. Really big fan of low cost API models, noone does it better than DS. Until VRAM price is low enough to run models locally, this is the way to go. The subsidized subscription model won't last, API pricing "feels" closer to a true sustainable business model.

Indeed. My fellow software engineers keep complaining about using up all their Claude tokens within an hour... Whilst I'll be rocking DS flash for the entire day. Sure it gets a few things wrong here and there, but that's when you pull out the Claude models or whatever for those tricky tasks.

what plan are your 'fellow software engineers' using? I have a hard time even using up the Fable part of my allowance in a week of coding.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#78
My problem with DS flash/pro is that they don’t push back on obvious bullshit, both irl and code [0] but it’s a great implementer workhorse if you give it _very_ detailed specs.

[0] https://petergpt.github.io/bullshit-benchmark/viewer/index.v...

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#79
post #51

Earlier quoted context omitted.

You’re correct. The hard line for me is ensuring that any scripture presented to the user is verified correct. LLMs can’t be trusted in this regard. One feature of the app is that all scripture is verified and what’s show to the user doesn’t come from the LLM at all and instead a trusted source. I think exploring scripture this way does not alleviate you from struggling to learn and apply it. It hasn’t for me.

I don't necessarily mean reguritating it, but choosing which part of the scripture to surface to the user is already some interpretation/choice. Even the devil can quote scripture (I'm playing the devil's advocate here).

Very true. The verses selected come from a tool call. The LLM queries the app for relevant verses. It picks the keywords - so perhaps there’s bias there but the verses are handed to the LLM.

But there are ways to control and constrain the LLMs and what the user is presented with.

These are all top of mind for me and why I felt there could be a better option than asking ChatGPT directly.

Re: DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

#80

Earlier quoted context omitted.

we need a benchmark website benchmark

A benchmark website to benchmark benchmark websites? Or a benchmark to benchmark benchmarks?

website that benchmarks benchmarks is benchmark benchmark website

benchmark website benchmark is indeed a benchmark that benchmarks websites with benchmarks (but it can be shown outside websites as well, it's not picky)

Post reply on HN