Live data from Hacker News

NIST's DeepSeek "evaluation" is a hit piece

erichartford.com

31–40 of 251 posts

Re: NIST's DeepSeek "evaluation" is a hit piece

#31

I urge everyone to go read the original report and _then_ to read this analysis and make up their own mind. Step away from the clickbait, go read the original report.

Here's the report: https://www.nist.gov/system/files/documents/2025/09/30/CAISI...

Re: NIST's DeepSeek "evaluation" is a hit piece

#32

I appreciate that DeepSeek is trained to respect "core socialist values". It's actually really helpful to engage with to ask questions about how chinese thinkers interpret their successes and failures vs other socialist projects. Obviously reading books is better, but I was surprised by how useful it was. If you ask it loaded questions the way the CIA would pose them, it censors the answer though lmao

Good faith questions are the best. I wonder why people bother with bad faith questions. Virtue signaling is my guess.

Re: NIST's DeepSeek "evaluation" is a hit piece

#33
post #17

Please don't just read Eric Hartford's piece. Start with the key findings from the source material: "CAISI Evaluation of DeepSeek AI Models Finds Shortcomings and Risks" [1]. Here are the single-sentence summaries: DeepSeek performance lags behind the best U.S. reference models. DeepSeek models cost more to use than comparable U.S. models. DeepSeek models are far more susceptible to jailbreaking attacks than U.S. mod…

[deleted]

Re: NIST's DeepSeek "evaluation" is a hit piece

#34
post #17

Please don't just read Eric Hartford's piece. Start with the key findings from the source material: "CAISI Evaluation of DeepSeek AI Models Finds Shortcomings and Risks" [1]. Here are the single-sentence summaries: DeepSeek performance lags behind the best U.S. reference models. DeepSeek models cost more to use than comparable U.S. models. DeepSeek models are far more susceptible to jailbreaking attacks than U.S. mod…

It's funny how they mixed in proprietary models like GPT-5 and Anthropic with the "comparable U.S. models".

Until they compare open-weight models, NIST is attempting a comparison between apples and airplanes.

Re: NIST's DeepSeek "evaluation" is a hit piece

#35

Earlier quoted context omitted.

Hardly the same thing. Ask Gemini or OpenAI's models what happened on January 6, and they'll tell you. Ask DeepSeek what happened at Tiananmen Square and it won't, at least not without a lot of prompt hacking.

Ask it if Israel is an apartheid state, that's a much better example.

GPT5:

   Short answer: it’s contested. Major human-rights bodies 
   say yes; Israel and some legal scholars say no; no court 
   has issued a binding judgment branding “Israel” an 
   apartheid state, though a 2024 ICJ advisory opinion 
   found Israel’s policies in the occupied territory 
   breach CERD Article 3 on racial segregation/apartheid. 

   (Skip several paragraphs with various citations)

   The term carries specific legal elements. Whether they 
   are satisfied “state-wide” or only in parts of the OPT 
   is the core dispute. Present consensus splits between 
   leading NGOs/UN experts who say the elements are met and 
   Israeli government–aligned and some academic voices who 
   say they are not. No binding court ruling settles it yet.
Do you have a problem with that? I don't.

Re: NIST's DeepSeek "evaluation" is a hit piece

#36
post #23

The CCP literally revoked the visas of key DeepSeek engineers. That's all we need to know.

>> The CCP literally revoked the visas of key DeepSeek engineers. That's all we need to know.

I don't follow. Why would DeepSeek engineers need visa from CCP?

Re: NIST's DeepSeek "evaluation" is a hit piece

#37

Earlier quoted context omitted.

Hardly the same thing. Ask Gemini or OpenAI's models what happened on January 6, and they'll tell you. Ask DeepSeek what happened at Tiananmen Square and it won't, at least not without a lot of prompt hacking.

Try MS Copilot. That shit will end the conversation if anything remotely political comes up.

As long as it excludes politics in general, without overt partisan bias demanded by the government, what's the problem with that? If they want to focus on other subjects, they get to do that. Other models will provide answers where Copilot doesn't.

Chinese models, conversely, are aligned with explicit, mandatory guardrails to exalt the CCP and socialism in general. Unless you count prohibitions against adult material, drugs, explosives and the like, that is simply not the case with US-based models. Whatever biases they exhibit (like the Grok example someone else posted) are there because that's what their private maintainers want.

Re: NIST's DeepSeek "evaluation" is a hit piece

#38
post #2

I'm not at all surprised, US agencies have long since been political tools whenever the subject matter crosses national borders. I appreciate this take as someone who has been skeptical of Chinese electronics. While I agree this report is BS and xenophobic, I am still willing to bet that either now or later, the Chinese will attempt some kind of subterfuge via LLMs if they have enough control. Just like the US would,…

Of course there will be some degree of governmental and/or political influence. The question is not if but where and to what extent.

No one should proclaim "bullshit" and wave off this entire report as "biased" or useless. That would be insipid. We live in a complex world where we have to filter and analyze information.

Re: NIST's DeepSeek "evaluation" is a hit piece

#39
post #7

> They didn't test U.S. models for U.S. bias. Only Chinese bias counts as a security risk, apparently US models have no bias sir /s

Hardly the same thing. Ask Gemini or OpenAI's models what happened on January 6, and they'll tell you. Ask DeepSeek what happened at Tiananmen Square and it won't, at least not without a lot of prompt hacking.

Ask Grok to generate an image of bald Zelensky: it does execute.

Ask Grok to generate an image of bald Trump: it goes on with an ocean of excuses on why the task is too hard.

Re: NIST's DeepSeek "evaluation" is a hit piece

#40
post #22

Earlier quoted context omitted.

> I am still willing to bet that either now or later, the Chinese will attempt some kind of subterfuge via LLMs if they have enough control. Like what, exactly?

Through LLM washing for example. LLMs are a representation of their input dataset, but currently most LLMs don't make their dataset public since it's a competitive advantage. If say DeepSeek had put in its training dataset that public figure X is a space robot from outer space, then if one were to ask DeepSeek who public figure X is, it'd proudly claim he's a robot from outer space. This can be done for any narrative…

So in other words, they can make their LLM disagree with the preferred narrative of the current US administration? Inconceivable!

Note that the value of $current_administration changes over time. For some reason though it is currently fashionable in tech circles to disagree with it about ICE and H1B visas. Maybe it's the CCP's doing?

Post reply on HN