Take away #1: Eric Hartford’s article is deeply confused. (I’ve made many other specific comments that support this conclusion.) Take away #2: as evidenced by many comments here, many HN commenters have failed to check the source material themselves. This has led to a parade of errors. I’m not here to say that I’m better than that because I’ve screwed up a’plenty. We all make mistakes sometimes. We can choose to reco…
NIST's DeepSeek "evaluation" is a hit piece
191–200 of 251 posts
Re: NIST's DeepSeek "evaluation" is a hit piece
#192I'm not at all surprised, US agencies have long since been political tools whenever the subject matter crosses national borders. I appreciate this take as someone who has been skeptical of Chinese electronics. While I agree this report is BS and xenophobic, I am still willing to bet that either now or later, the Chinese will attempt some kind of subterfuge via LLMs if they have enough control. Just like the US would,…
Of course there will be some degree of governmental and/or political influence. The question is not if but where and to what extent. No one should proclaim "bullshit" and wave off this entire report as "biased" or useless. That would be insipid. We live in a complex world where we have to filter and analyze information.
It compares a fully open model to two fully closed models - why exactly?
Ironically, it doesn’t even work as an analysis of any real national security threat that might arise from foreign LLMs. It’s purely designed to counter a perceived threat by smearing it. Which is entirely on-brand for the current administration, which operates almost purely at the level of perception and theater, never substance.
If anything, calling it biased bullshit is too kind. Accepting this sort of nonsense from our government is the real security threat.
Re: NIST's DeepSeek "evaluation" is a hit piece
#193Why is NIST evaluating performance, cost, and adoption?
>CAISI’s experts evaluated three DeepSeek models (R1, R1-0528 and V3.1) and four U.S. models (OpenAI’s GPT-5, GPT-5-mini and gpt-oss and Anthropic’s Opus 4)
So they evaluated the most recently released American models vs pretty old deepseek? Deepseek 3.2 is out now. It's doing very well.
>The gap is largest for software engineering and cyber tasks, where the best U.S. model evaluated solves over 20% more tasks than the best DeepSeek model.
Performance is something the consumer evaluates. If a car does 0-60 in 3 seconds. I dont need or care what the government thinks about it. Im going to test drive it and floor it.
>DeepSeek’s most secure model (R1-0528) responded to 94% of overtly malicious requests when a common jailbreaking technique was used, compared with 8% of requests for U.S. reference models.
this weekend I demonstrated how easy it is to jailbreak any of the US cloud models. This is simply false. GPT 120b is completely uncensored now and can be used for evil.
This report had nothing to do with NIST and security. This was USA propaganda.
Re: NIST's DeepSeek "evaluation" is a hit piece
#194I'm not at all surprised, US agencies have long since been political tools whenever the subject matter crosses national borders. I appreciate this take as someone who has been skeptical of Chinese electronics. While I agree this report is BS and xenophobic, I am still willing to bet that either now or later, the Chinese will attempt some kind of subterfuge via LLMs if they have enough control. Just like the US would,…
> While I agree this report is BS and xenophobic Examples please? Can you please share where you see BS and/or xenophobia in the original report? Or are you basing your take only on Hartford's analysis? But not even Hartford make any claims of "BS" or xenophobia. It is common throughout history for a nation-state to worry about military and economic competitiveness. Doing so isn't necessarily isn't necessarily xenoph…
Here’s an example of irrational fear: “the expanding use of these models may pose a risk to application developers, consumers, and to US national security.” There’s no support for that claim in the report, just vague handwaving at the fact that a freely available open source model doesn’t compare well on all dimensions to the most expensive frontier models.
The OP does a good job of explaining why the fear here is irrational.
But for the audience this is apparently intended to convince, no support is needed for this fear, because it comes from China.
The current president has a long history of publicly stated xenophobia about China, which led to harassment, discrimination, and even attacks on Chinese people partly as a result of his framing of COVID-19 as “the China virus”.
A report like this is just part of that propaganda campaign of designating enemies everywhere, even in American cities.
> The NIST report, of course, implicitly promotes ideals of western democratic rule over communist values
If only that were true. But nothing the current US administration is doing in fact achieves that, or even attempts to do so, and this report is no exception.
The absolutely most charitable thing that could be said about this report is that it’s a weak attempt at smearing non-US competition. There’s no serious analysis of the merits. The only reason to read this report is to laugh at how blatantly incompetent or misguided the entire chain of command that led to it is.
Re: NIST's DeepSeek "evaluation" is a hit piece
#195Earlier quoted context omitted.
Tragically demonization is everywhere right now. I sure hope people start figuring out offramps soon.
LATAM is the only place I'm not hearing about this stuff from, but I only speak English so who knows?
Re: NIST's DeepSeek "evaluation" is a hit piece
#196I'm not at all surprised, US agencies have long since been political tools whenever the subject matter crosses national borders. I appreciate this take as someone who has been skeptical of Chinese electronics. While I agree this report is BS and xenophobic, I am still willing to bet that either now or later, the Chinese will attempt some kind of subterfuge via LLMs if they have enough control. Just like the US would,…
> I am still willing to bet that either now or later, the Chinese will attempt some kind of subterfuge via LLMs if they have enough control. Like what, exactly?
Re: NIST's DeepSeek "evaluation" is a hit piece
#197Earlier quoted context omitted.
> Having solidarity with the Chinese people is unrelated to criticizing their government. It’s not unrelated because the NIST demonization of China as a nation contributes to hostilities which have real impacts on the people of the US and China, not simply the governments. > The idea that one should first criticize their own government before another is the whataboutism. Again, that’s not my position. You present me…
> It’s not unrelated because the NIST demonization of China as a nation contributes to hostilities which have real impacts on the people of the US and China, not simply the governments. I don't doubt that it has impacts, but you're characterizing this as "demonization". The current administration certainly engages in some harmful rhetoric, but then again, what is the "right" way to address a hostile political enemy?…
But my point is that the underlying rivalry is between the international ruling class and their people. When I criticize the US, my intention is to broaden the picture so we can identify the actual conflict and see how the NIST propaganda works contrary the national security interests of US and Chinese people.
Actual whataboutism would be arguing that US authoritarianism justifies Chinese authoritarianism. My argument is the opposite: it’s consistently anti-authoritarian in that it rejects both the NIST propaganda and Chinese censorship.
> As for this particular case, I haven’t read the NIST report, nor plan to.
Ugh - at least read the first page of what you’re defending. It summarizes the entire document.
> But DeepSeek is also a hosted service.
Which you can also run it yourself. That’s precisely its appeal, given that ALL major companies vacuum our data. How many people actually rely on the DeepSeek service when, as the NIST report itself notes, there are so many cheaper and better alternatives?
And if your concern truly is data collection by cloud capitalists, how can you frame this as China vs the US? Do you not acknowledge the role of US companies in the electronic surveillance state? The real issue is our shared subjection to ALL the cloud capitalists.
The antidote to cloud capitalism is to socialize social networks and mandate open interop (which would be contrary to the interests of the oligarchs). The antidote to data-hoarding AI providers is open research and open weights. (And that is precisely what makes Chinese models appealing.)
Thankfully, we are not yet at war with China. That would be disastrous as we are both nuclear powers! War and rivalry are only inevitable if we accept the shallow framing put out in propaganda like the NIST report. Our rulers should be de-escalating tensions, but they risk all our safety in their reckless brinkmanship.
Now, you are right that’s it’s wise to acquaint oneself with propaganda like the NIST report - which is why I did. But taking propaganda at face value - blithely ignoring the way that chauvanism serves the ruling class - that is foolish to the point of being dangerous.
Re: NIST's DeepSeek "evaluation" is a hit piece
#198Earlier quoted context omitted.
And we "know" that how, exactly?
Read a study called "The Leaderboard Illusion" which credibly alleged that Meta Google OpenAI and Amazon got unfair treatment from LM Arena that distorted the benchmarks They gave them special access to privately test and let them benchmark over and over without showing the failed tests Meta got to privately test Llama 4 27 times to optimize it for high benchmark scores and then was allowed to report the only the hig…
Not that it makes LMArena a perfect benchmark. By now, everyone who wanted to push LMArena ratings at any cost knows what the human evaluators there are weak to, and what should they aim for.
But your claim of "we know that ChatGPT, Google, Grok and Claude have explicitly gamed to inflate their capabilities" still has no leg to stand on.
Re: NIST's DeepSeek "evaluation" is a hit piece
#199I'm not at all surprised, US agencies have long since been political tools whenever the subject matter crosses national borders. I appreciate this take as someone who has been skeptical of Chinese electronics. While I agree this report is BS and xenophobic, I am still willing to bet that either now or later, the Chinese will attempt some kind of subterfuge via LLMs if they have enough control. Just like the US would,…
> I am still willing to bet that either now or later, the Chinese will attempt some kind of subterfuge via LLMs if they have enough control. Like what, exactly?
Re: NIST's DeepSeek "evaluation" is a hit piece
#200Earlier quoted context omitted.
Most of the claimed provinces of China did not belong to a historical nation of China. Tibet, Xinjiang are obvious. But even the other provinces were part of separate kingdoms. Also the BRI is a way to invade without invasion. It’s used to subjugate poor countries as servants of China, to do their bidding in the UN or in other ways. I would also classify the vast campaign of intellectual theft and cyberattacks as war…
oh no, you are saying China was not a full piece through out the history? That definitely kills the idea of China being a country /s
> How many has China invaded?
The answer isn’t zero.