Live data from Hacker News

NIST's DeepSeek "evaluation" is a hit piece

erichartford.com

81–90 of 251 posts

Re: NIST's DeepSeek "evaluation" is a hit piece

#81
post #74

Earlier quoted context omitted.

What about a rational distaste for the CCP?

Not sure how it’s rational if you don’t extend the same distaste to our authoritarian government. Concentration camps, genocide, suppressing free speech, suspending due process. That’s what it’s up to these days. To say nothing of the effectively dictatorial control the ultra wealthy have over public policy. Sinophobia is a distraction from our problems at home. That’s its purpose.

That's whataboutism at its purest. It's perfectly possible to criticize any government, whether your own or foreign.

Claiming that every criticism is tantamount to racism is what's distracting from discussing actual problems.

Re: NIST's DeepSeek "evaluation" is a hit piece

#82

I urge everyone to go read the original report and _then_ to read this analysis and make up their own mind. Step away from the clickbait, go read the original report.

Here's the report: https://www.nist.gov/system/files/documents/2025/09/30/CAISI...

> DeepSeek models cost more to use than comparable U.S. models

They compare DeepSeek v3.1 to GPT-5 mini. Those have very different sizes, which makes it a weird choice. I would expect a comparison with GPT-5 High, which would likely have had the opposite finding, given the high cost of GPT-5 High, and relatively similar results.

Granted, DeepSeek typically focuses on a single model at a time, instead of OpenAI's approach to a suite of models of varying costs. So there is no model similar to GPT-5 mini, unlike Alibaba which has Qwen 30B A3B. Still, weird choice.

Besides, DeepSeek has shown with 3.2 that it can cut prices in half through further fundamental research.

Re: NIST's DeepSeek "evaluation" is a hit piece

#83

Earlier quoted context omitted.

> I am still willing to bet that either now or later, the Chinese will attempt some kind of subterfuge via LLMs if they have enough control. Like what, exactly?

Like generating vulnerable code given a specific prompt/context. I also don't think it's just China, the US will absolutely order American providers to do the same. It's a perfect access point for installing backdoors into foreign systems.

> Like generating vulnerable code given a specific prompt/context.

That's easy (well, possible) to detect. I'd go the opposite way - sift the code that is submitted to identify espionage targets. One example: if someone submits a piece of commercial code that's got a vulnerability, you can target previous versions of that codebase.

I'd be amazed if that wasn't happening already.

Re: NIST's DeepSeek "evaluation" is a hit piece

#84

Earlier quoted context omitted.

Up until recently, I would have reminded you that the US government (admittedly unlike the Chinese government) has no legal authority to order anybody to do anything like that. Not only that, but if it asked , it'd be well advised to ask nicely, because it also has no legal authority to demand that anybody keep such a request secret. And no, evil as it is, the "National Security Letter" power doesn't in fact cover an…

> Up until recently, I would have reminded you that the US government (admittedly unlike the Chinese government) has no legal authority to order anybody to do anything like that. I'm not sure how closely you've been following, but the US government has a long history of doing things they don't have legal authority to do.

Why would you need legal authority when you have whole host of legal tools you can use. Making life a difficult for anyone or any company is simple enough. Just by state finally doing their job properly for example.

Re: NIST's DeepSeek "evaluation" is a hit piece

#85

Earlier quoted context omitted.

Hardly the same thing. Ask Gemini or OpenAI's models what happened on January 6, and they'll tell you. Ask DeepSeek what happened at Tiananmen Square and it won't, at least not without a lot of prompt hacking.

Ask Grok to generate an image of bald Zelensky: it does execute. Ask Grok to generate an image of bald Trump: it goes on with an ocean of excuses on why the task is too hard.

FWIW, I can't reproduce this example - it generates both images fine: https://ibb.co/NdYx1R4p

Re: NIST's DeepSeek "evaluation" is a hit piece

#86

Earlier quoted context omitted.

GPT5: Short answer: it’s contested. Major human-rights bodies say yes; Israel and some legal scholars say no; no court has issued a binding judgment branding “Israel” an apartheid state, though a 2024 ICJ advisory opinion found Israel’s policies in the occupied territory breach CERD Article 3 on racial segregation/apartheid. (Skip several paragraphs with various citations) The term carries specific legal elements. Wh…

I better not poke that hornets nest any further, but yeah I made my point.

I better not poke that hornets nest any further, but yeah I made my point.

Yes, I can certainly see why you wouldn't want to go any further with the conversation.

Re: NIST's DeepSeek "evaluation" is a hit piece

#87
post #60

Insightful post, thanks for sharing. What are people's experiences with the uncensored Dolphin model the author has made?

> What are people's experiences with the uncensored Dolphin model the author has made?

My take? The best way to know is to build your own eval framework and try it yourself. The "second best" way would be to find someone else's eval which is sufficiently close to yours. (But how would you know if another's eval is close enough if you haven't built your own eval?)

Besides, I wouldn't put much weight on a random commenter here. Based on my experiences on HN, I highly discount what people say because I'm looking for clarity, reasoning, and nuance. My discounting is 10X worse for ML or AI topics. People seem too hurried, jaded, scarred, and tribal to seek the truth carefully, so conversations are often low quality.

So why am I here? Despite all the above, I want to participate in and promote good discussion. I want to learn and to promote substantive discussion in this community. But sometimes it feels like this: https://xkcd.com/386/

Re: NIST's DeepSeek "evaluation" is a hit piece

#88

Earlier quoted context omitted.

Ask Grok to generate an image of bald Zelensky: it does execute. Ask Grok to generate an image of bald Trump: it goes on with an ocean of excuses on why the task is too hard.

FWIW, I can't reproduce this example - it generates both images fine: https://ibb.co/NdYx1R4p

I asked it in french a few days back and it went on explaining me how hard this would be. Thanks for the update.

EDIT: I tried it right now and it did generate the image. I don't know what happened then...

Re: NIST's DeepSeek "evaluation" is a hit piece

#89
post #83

Earlier quoted context omitted.

Like generating vulnerable code given a specific prompt/context. I also don't think it's just China, the US will absolutely order American providers to do the same. It's a perfect access point for installing backdoors into foreign systems.

> Like generating vulnerable code given a specific prompt/context. That's easy (well, possible) to detect. I'd go the opposite way - sift the code that is submitted to identify espionage targets. One example: if someone submits a piece of commercial code that's got a vulnerability, you can target previous versions of that codebase. I'd be amazed if that wasn't happening already.

The thing with chinese models for the most part is that they are open weights so it depends on if somebody is using their api or not.

Sure, maybe something like this can happen if you use the deepseek api directly which could have chinese servers but that is a really long strech but to give the benefit of doubt, maybe

but your point becomes moot if somebody is hosting their own models. I have heard glm 4.6 is really good comparable to sonnet and can definitely be used as a cheaper model for some stuff, currently I think that the best way might be to use something like claude 4 or gpt 5 codex or something to generate a detailed plan and then execute it using the glm 4.6 model preferably by using american datacenter providers if you are worried about chinese models without really worrying about atleast this tangent and getting things done at a cheaper cost too

Re: NIST's DeepSeek "evaluation" is a hit piece

#90
Considering DeepSeek had a peer-reviewd analysis in nature https://www.nature.com/articles/s41586-025-09422-z relaes just last month with indipendent researcher affriming that the open model has some issues(acknowldged in the writeup) , well inclined to agree with the articles author , the NIST evaluation looks more like a politcal hatchet job with a bit of projection going on(ala this is what the US would do if they were in that position). To be fair the paranoia has a basis in that whenever there is tech-leverage the US TLA subverts it for espionage like the CryptoAG episode. Or recently the whole hoopla about Huawei in the EU , which after relentless searches only turned up bad coding practices rather than anything malicious. At this pint it would be better for the whole field that these models exist as well as Kimi, Qwen etc as the downward pressure on cost/capabilities leads to commoditisation and the whole race to build a ecogeopolitical moat goes away.
Post reply on HN