Live data from Hacker News

NIST's DeepSeek "evaluation" is a hit piece

erichartford.com

211–220 of 251 posts

Re: NIST's DeepSeek "evaluation" is a hit piece

#211
post #162

Earlier quoted context omitted.

did you even read the article? when you download open source deep seek model and run it yourself - zero packets are being transmitted. thereby disproving the fundamental claim in the NIST report (additionally NIST doesn't provide any evidence to support their claim) This is basic science and no amount of politicking should ever challenge something this fundamental!

> did you even read the article? I am not going to dignify this with a response. Please review the hacker news guidelines.

Notice the similarity between the above comment and what the HN Guidelines [1] advise against:

> Please don't comment on whether someone read an article. "Did you even read the article? It mentions that" can be shortened to "The article mentions that".

[1]: https://news.ycombinator.com/newsguidelines.html

Re: NIST's DeepSeek "evaluation" is a hit piece

#212

Earlier quoted context omitted.

did you even read the article? when you download open source deep seek model and run it yourself - zero packets are being transmitted. thereby disproving the fundamental claim in the NIST report (additionally NIST doesn't provide any evidence to support their claim) This is basic science and no amount of politicking should ever challenge something this fundamental!

Meanwhile we know for a fact that Gemini for example uses chat logs to build a social graph and Google is complicit in NSA surveillance Not to mention Anthropic says Claude will eventually automatically report you to authorities if you ask it to do something "unethical"

> Not to mention Anthropic says Claude will eventually automatically report you to authorities if you ask it to do something "unethical"

Are you referring to the situation described in the May 22, 2025 article by Carl Franzen in VentureBeat [1]? If so, at a minimum, one should recognize the situation is complex enough to warrant a careful look for yourself to wade through the confusion. Speaking for myself, I don't have anything close to a "final take" yet.

[1]: https://venturebeat.com/ai/anthropic-faces-backlash-to-claud...

Re: NIST's DeepSeek "evaluation" is a hit piece

#213

Earlier quoted context omitted.

Read a study called "The Leaderboard Illusion" which credibly alleged that Meta Google OpenAI and Amazon got unfair treatment from LM Arena that distorted the benchmarks They gave them special access to privately test and let them benchmark over and over without showing the failed tests Meta got to privately test Llama 4 27 times to optimize it for high benchmark scores and then was allowed to report the only the hig…

Which is one study that touches exactly one benchmark - and "credibly alleged" is being way too generous to it. The only case that was anywhere close to being proven LMArena fraud is Meta and Llama 4. Which is a nonentity now - nowhere near SOTA on anything, LMArena included. Not that it makes LMArena a perfect benchmark. By now, everyone who wanted to push LMArena ratings at any cost knows what the human evaluators…

There are a lot of other cases that extend well beyond LMArena where it was shown certain benchmark performance increases by the major US labs were only attributable to being over-optimized for the specific benchmarks. Some in ways that are not explainable by the benchmark tests merely contaminating the corpus.

There are cases where merely rewording the questions or assigning different letters to the answer dropped models like Llama 30% in the evaluations while others were unchanged

Open-LLM-Leaderboard had to rate limit because a "handful of labs" were doing so many evals in a single day that it hogged the entire eval cluster

“Coding Benchmarks Are Already Contaminated” (Ortiz et al., 2025) “GSM-PLUS: A Re-translation Reveals Data Contamination” (Shi et al., ACL 2024). “Prompt-Tuning Can Add 30 Points to TruthfulQA” (Perez et al., 2023). “HellaSwag Can Be Gamed by a Linear Probe” (Rajpurohit & Berg-Kirkpatrick, EMNLP 2024). “Label Bias Explains MMLU Jumps” (Hassan et al., arXiv 2025) “HumanEval-Revival: A Re-typed Test for LLM Coding Ability” (Yang & Liu, ICML 2024 workshop). “Data Contamination or Over-fitting? Detecting MMLU Memorisation in Open LLMs” (IBM, 2024)

And yes I relied on LLM to summarize these instead of reading the full papers

Re: NIST's DeepSeek "evaluation" is a hit piece

#214
post #166

Take away #1: Eric Hartford’s article is deeply confused. (I’ve made many other specific comments that support this conclusion.) Take away #2: as evidenced by many comments here, many HN commenters have failed to check the source material themselves. This has led to a parade of errors. I’m not here to say that I’m better than that because I’ve screwed up a’plenty. We all make mistakes sometimes. We can choose to reco…

Since you read this report in full, can you please give me the authors' names? I did read it (partially) and didn't find any name. I am, for sure, a beginner in NIST's reports reading, but I found out that a lot of NIST reports are signed (ie you can know who wrote the report, on behalf of whom if external contractor).

> can you please give me the authors' names?

The names of the author(s) are not given.

Re: NIST's DeepSeek "evaluation" is a hit piece

#215
post #78
post #63

Earlier quoted context omitted.

And what are the downsides?

Making the interests of a population subservient to those of a foreign state. Now if that sounds nice to you please, by all means, do just migrate to China.

It was somewhat sarcastic comment inviting reader to replace China with the US and the US with Russia or the Ukraine.

China doesn't offer citizenship for foreigners but if I wanted to see the cities of the future I could go there visa-free.

Re: NIST's DeepSeek "evaluation" is a hit piece

#216

Earlier quoted context omitted.

Is this some kind of satire or are you just completely ignorant of European/US history? Either way its laughable to even compare IP theft to the invasion of Iraq or bombing of Cambodia. How do you think the industrial revolution got started in the US, they just did it on their own? Not to mention that the entire US was stolen from the natives.

No, it’s not laughable. Your insinuation that China doesn’t invade other countries, meant to imply they haven’t engaged in warfare, was false. And yes IP theft is comparable to invasions and often worse. > Not to mention that the entire US was stolen from the natives. This is partially true. But partially false. You can figure out why if you’re curious.

> And yes IP theft is comparable to invasions and often worse.

This assertion smells more American than a Big Mac. Do you have any actual citations?

In a free market, lowering the barrier-to-entry in a given market tends to increase competition. Industry-scale IP theft really only damages your economy if the rent-seekers rely on low competition. A country with a strong primary/secondary sector (resources and manufacturing) never needs to rely on protecting precious IP. America has already lost if we depend on playing keep-away with F-35 schematics for basic doctrinal advantage.

Re: NIST's DeepSeek "evaluation" is a hit piece

#217

Earlier quoted context omitted.

No, it’s not laughable. Your insinuation that China doesn’t invade other countries, meant to imply they haven’t engaged in warfare, was false. And yes IP theft is comparable to invasions and often worse. > Not to mention that the entire US was stolen from the natives. This is partially true. But partially false. You can figure out why if you’re curious.

> And yes IP theft is comparable to invasions and often worse. This assertion smells more American than a Big Mac. Do you have any actual citations? In a free market, lowering the barrier-to-entry in a given market tends to increase competition. Industry-scale IP theft really only damages your economy if the rent-seekers rely on low competition. A country with a strong primary/secondary sector (resources and manufact…

All of that is just a wild justification for large-scale economic damage to another country. In other words, warfare.

Re: NIST's DeepSeek "evaluation" is a hit piece

#218

Earlier quoted context omitted.

> And yes IP theft is comparable to invasions and often worse. This assertion smells more American than a Big Mac. Do you have any actual citations? In a free market, lowering the barrier-to-entry in a given market tends to increase competition. Industry-scale IP theft really only damages your economy if the rent-seekers rely on low competition. A country with a strong primary/secondary sector (resources and manufact…

All of that is just a wild justification for large-scale economic damage to another country. In other words, warfare.

Hybrid warfare. Go bomb China for Salt Typhoon if it makes you feel any better, they still have the upper hand. Obsessing over retaliation instead of defense is precisely what China wants to provoke, it manufactures global consent to destroy America. No nation wants to coexist with a hegemon that goes nuclear whenever they're outdone.

When we forego obvious solutions ("hmm maybe telecoms need to be held to higher standards") and jump to war, America forfeits the competitive advantage and exacerbates the issue. For all of China's authoritarian misgivings, this is how they win.

Re: NIST's DeepSeek "evaluation" is a hit piece

#219
post #199

Earlier quoted context omitted.

> I am still willing to bet that either now or later, the Chinese will attempt some kind of subterfuge via LLMs if they have enough control. Like what, exactly?

The open source models are already heavily censored in ways the CCP likes, such as pretending the Tianamen Square massacre never happened. I expect they will go the TikTok route and crank that up to 11 over time, promoting topics that are divisive to the the US (and other adversaries) and outputting heavily biased results in ways that range from subtle to blatant.

That’s not heavy censorship. That’s a bit of censorship.

Re: NIST's DeepSeek "evaluation" is a hit piece

#220

Earlier quoted context omitted.

Honestly, I think this article is itself the hit piece (against NIST or America). And it is the one with inflammatory language.

Isn’t America currently killing its citizens with its own military? I would trust them even less now.

They're not, and I think you should trust whoever told you that even less now.
Post reply on HN