I urge everyone to go read the original report and _then_ to read this analysis and make up their own mind. Step away from the clickbait, go read the original report.
NIST's DeepSeek "evaluation" is a hit piece
31–40 of 251 posts
Re: NIST's DeepSeek "evaluation" is a hit piece
#32I appreciate that DeepSeek is trained to respect "core socialist values". It's actually really helpful to engage with to ask questions about how chinese thinkers interpret their successes and failures vs other socialist projects. Obviously reading books is better, but I was surprised by how useful it was. If you ask it loaded questions the way the CIA would pose them, it censors the answer though lmao
Re: NIST's DeepSeek "evaluation" is a hit piece
#33Please don't just read Eric Hartford's piece. Start with the key findings from the source material: "CAISI Evaluation of DeepSeek AI Models Finds Shortcomings and Risks" [1]. Here are the single-sentence summaries: DeepSeek performance lags behind the best U.S. reference models. DeepSeek models cost more to use than comparable U.S. models. DeepSeek models are far more susceptible to jailbreaking attacks than U.S. mod…
Re: NIST's DeepSeek "evaluation" is a hit piece
#34Please don't just read Eric Hartford's piece. Start with the key findings from the source material: "CAISI Evaluation of DeepSeek AI Models Finds Shortcomings and Risks" [1]. Here are the single-sentence summaries: DeepSeek performance lags behind the best U.S. reference models. DeepSeek models cost more to use than comparable U.S. models. DeepSeek models are far more susceptible to jailbreaking attacks than U.S. mod…
Until they compare open-weight models, NIST is attempting a comparison between apples and airplanes.
Re: NIST's DeepSeek "evaluation" is a hit piece
#35Earlier quoted context omitted.
Hardly the same thing. Ask Gemini or OpenAI's models what happened on January 6, and they'll tell you. Ask DeepSeek what happened at Tiananmen Square and it won't, at least not without a lot of prompt hacking.
Ask it if Israel is an apartheid state, that's a much better example.
Short answer: it’s contested. Major human-rights bodies
say yes; Israel and some legal scholars say no; no court
has issued a binding judgment branding “Israel” an
apartheid state, though a 2024 ICJ advisory opinion
found Israel’s policies in the occupied territory
breach CERD Article 3 on racial segregation/apartheid.
(Skip several paragraphs with various citations)
The term carries specific legal elements. Whether they
are satisfied “state-wide” or only in parts of the OPT
is the core dispute. Present consensus splits between
leading NGOs/UN experts who say the elements are met and
Israeli government–aligned and some academic voices who
say they are not. No binding court ruling settles it yet.
Do you have a problem with that? I don't.Re: NIST's DeepSeek "evaluation" is a hit piece
#36The CCP literally revoked the visas of key DeepSeek engineers. That's all we need to know.
I don't follow. Why would DeepSeek engineers need visa from CCP?
Re: NIST's DeepSeek "evaluation" is a hit piece
#37Earlier quoted context omitted.
Hardly the same thing. Ask Gemini or OpenAI's models what happened on January 6, and they'll tell you. Ask DeepSeek what happened at Tiananmen Square and it won't, at least not without a lot of prompt hacking.
Try MS Copilot. That shit will end the conversation if anything remotely political comes up.
Chinese models, conversely, are aligned with explicit, mandatory guardrails to exalt the CCP and socialism in general. Unless you count prohibitions against adult material, drugs, explosives and the like, that is simply not the case with US-based models. Whatever biases they exhibit (like the Grok example someone else posted) are there because that's what their private maintainers want.
Re: NIST's DeepSeek "evaluation" is a hit piece
#38I'm not at all surprised, US agencies have long since been political tools whenever the subject matter crosses national borders. I appreciate this take as someone who has been skeptical of Chinese electronics. While I agree this report is BS and xenophobic, I am still willing to bet that either now or later, the Chinese will attempt some kind of subterfuge via LLMs if they have enough control. Just like the US would,…
No one should proclaim "bullshit" and wave off this entire report as "biased" or useless. That would be insipid. We live in a complex world where we have to filter and analyze information.
Re: NIST's DeepSeek "evaluation" is a hit piece
#39> They didn't test U.S. models for U.S. bias. Only Chinese bias counts as a security risk, apparently US models have no bias sir /s
Hardly the same thing. Ask Gemini or OpenAI's models what happened on January 6, and they'll tell you. Ask DeepSeek what happened at Tiananmen Square and it won't, at least not without a lot of prompt hacking.
Ask Grok to generate an image of bald Trump: it goes on with an ocean of excuses on why the task is too hard.
Re: NIST's DeepSeek "evaluation" is a hit piece
#40Earlier quoted context omitted.
> I am still willing to bet that either now or later, the Chinese will attempt some kind of subterfuge via LLMs if they have enough control. Like what, exactly?
Through LLM washing for example. LLMs are a representation of their input dataset, but currently most LLMs don't make their dataset public since it's a competitive advantage. If say DeepSeek had put in its training dataset that public figure X is a space robot from outer space, then if one were to ask DeepSeek who public figure X is, it'd proudly claim he's a robot from outer space. This can be done for any narrative…
Note that the value of $current_administration changes over time. For some reason though it is currently fashionable in tech circles to disagree with it about ICE and H1B visas. Maybe it's the CCP's doing?