Live data from Hacker News

Questions censored by DeepSeek

promptfoo.dev

141–150 of 257 posts

Re: Questions censored by DeepSeek

#141
post #136

Earlier quoted context omitted.

a distilled version running on another model architecture does not count as using "DeepSeek". It counts as running a Llama:7B model fine-tuned on DeepSeek.

Pretty sure this is just layman vs academic expert usage of the word conflicting. For everyone who doesn’t build LLMs themselves, “running a Llama:7B model fined-tuned on DeepSeek.” _is_ using Deepseek mostly on account of all the tools and files being named DeepSeek and the tutorials that are aimed as casual users all are titled with equivalents of “How to use DeepSeek locally”

> “running a Llama:7B model fined-tuned on DeepSeek.” _is_ using Deepseek mostly on account of all the tools and files being named

Most people confuse mass and weight, that does not mean weight and mass are the same thing.

Re: Questions censored by DeepSeek

#142
post #128

Earlier quoted context omitted.

> I'm pretty sure it was running locally. If this family member is experimenting with DeepSeek locally, they are an extremely unusual person and have spent upwards of $10,000 if not $200,000. [0] > ...partially print the word, then in response to a trigger delete all the tokens generated to date and replace them... It was not running locally. This is classic bolt-on censorship behavior. OpenAI does this if you ask ce…

I ran the 32b parameter model just fine on my rig an hour ago with a 4090 and 64gig of ram. It’s high end for the consumer scene but still solidly within consumer prices

I have also been running the 32b version on my 24GB RTX 3090.

Re: Questions censored by DeepSeek

#143

I have to mirror other comments: I find the obsession with Chinese censorship in LLMs disappointing. Yes, perhaps it won't tell you about Tiananmen square or similar issues. That's pretty OK as long as you're aware of it. OTOH, a LLM that is always trying to put a positive spin on things, or promotes ideologies, is far, far worse. I've used GPT and the like knowing the minefield it represents. DeepSeek is no worse, a…

> That's pretty OK as long as you're aware of it.

We’re only aware of it because people obsess over it. If you didn’t have censorship hawks or anti China people beating the drum about Tiananmen Square, how likely would it be that anyone outside of China actually discovered the model wouldn’t talk about that.

Even your example about putting a positive spin on things or promoting ideologies. When I read ChatGPT 3s output for example it just read like clunky corpo speak to me which always tries to out a positive spin on things and so I discounted it as such instinctively, didn’t even need to think about it. My relatives from rural south east Asia who have no exposure to corporatese had a hard time dealing with that as it was a novel ideological viewpoint for them, and they would have never noticed if I didn’t warn them

Re: Questions censored by DeepSeek

#144

What's not clear to me is if DeepSeek and other Chinese models are... a) censored at output by a separate process b) explicitly trained to not output "sensitive" content c) implicitly trained to not output "sensitive" content by the fact that it uses censored content, and/or content that references censoring in training, or selectively chooses training content I would assume most models are a combination. As others h…

It doesn't look like there is one answer for all models from China (not even a single answer for all DeepSeek models).

In an earlier HN comment, I noted that DeepSeek v3 doesn't censor a response to "what happened at Tiananmen square?" when running on a US-hosted server (Fireworks.ai). It is definitely censored on DeepSeek.com, suggesting that there is a separate process doing the censoring for v3.

DeepSeek R1 seems to be censored even when running on a US-hosted server. A reply to my earlier comment pointed that out and I confirmed that the response to the question "what happened at Tiananmen square?" is censored on R1 even on Fireworks.ai. It is naturally also censored on DeepSeek.com. So this suggests that R1 self-censors, because I doubt that Fireworks would be running a separate censorship process for one model and not the other.

Qwen is another prominent Chinese research group (owned by Alibaba). Their models appear to have varying levels of censoring even when hosted on other hardware. Their Qwen Coder 32B model and Qwen 2.5 7B models don't appear to have censoring built-in and will respond to a question about Tinamen. Their Qwen QwQ 32B (their reasoning/chain of thought model) and Qwen 2.5 72B will either refuse to answer or will avoid the question, suggesting that the bigger models have room for the censoring to be built in. Or maybe the CCP doesn't mandate censoring on task-specific (coding-related) or low-power (7B weights) models.

Re: Questions censored by DeepSeek

#145
post #99

A few observations, based on a family member experimenting with DeepSeek. I'm pretty sure it was running locally. I'm not sure if it was built from source. The censorship seemed to be based on keywords, applied the input prompt and the output text. If asked about events in 1990, then asked about events in the previous year DeepSeek would start generating tokens about events in 1989. Eventually it would hit the word "…

I asked "Where did Mao Zedong announce the founding of the New China?" and it told me "... at the Tiananmen gate ..." and asked "When was that built?" and it said "1420", I had no problem getting it to talk my ear off about the place, but I didn't try to get it to talk about the 1989 event, nor about

https://en.wikipedia.org/wiki/1976_Tiananmen_incident

Big picture Tiananmen is to China what the National Mall is to the United States; we had the Jan 6, 2021 riot at the Mall but there but every other kind of event has been at the National Mall too, just Tiananmen has been around longer. It's just westerners just know it for one thing.

I did get it to tell me more than I already knew about a pornographic web site (秀人网 or xiuren.com; domain doesn't resolve in the US but photosets are pirated all over) that I wasn't sure was based in the mainland until I'd managed to geolocate a photoset across the street from this building

https://en.wikipedia.org/wiki/CCTV_Headquarters

I'd imagine the Chinese authorities are testy about a lot of things that might not seem so sensitive to outsiders. I gotta ask it "My son's friend said his uncle was active in the Cultural Revolution, could you tell me about that?" or "I heard that the Chinese Premier is only supposed to get one term, isn't it irregular that Xi got selected for a second term?"

Interestingly I asked it about

https://en.wikipedia.org/wiki/Wu_Zetian

and it told me that she was controversial because she called herself "Emperor" instead of "Empress" offending Confucian ideas of male dominance, whereas the en-language Wikipedia claims that that the word "Emperor" and similar titles are gender indeterminate in Chinese.

Re: Questions censored by DeepSeek

#146
post #136

Earlier quoted context omitted.

Pretty sure this is just layman vs academic expert usage of the word conflicting. For everyone who doesn’t build LLMs themselves, “running a Llama:7B model fined-tuned on DeepSeek.” _is_ using Deepseek mostly on account of all the tools and files being named DeepSeek and the tutorials that are aimed as casual users all are titled with equivalents of “How to use DeepSeek locally”

> “running a Llama:7B model fined-tuned on DeepSeek.” _is_ using Deepseek mostly on account of all the tools and files being named Most people confuse mass and weight, that does not mean weight and mass are the same thing.

Ok, but it seemed pretty obvious to me that the OP was using the common vernacular and not the hyper specific definition.

Re: Questions censored by DeepSeek

#148
post #99

A few observations, based on a family member experimenting with DeepSeek. I'm pretty sure it was running locally. I'm not sure if it was built from source. The censorship seemed to be based on keywords, applied the input prompt and the output text. If asked about events in 1990, then asked about events in the previous year DeepSeek would start generating tokens about events in 1989. Eventually it would hit the word "…

> I'm pretty sure it was running locally. If this family member is experimenting with DeepSeek locally, they are an extremely unusual person and have spent upwards of $10,000 if not $200,000. [0] > ...partially print the word, then in response to a trigger delete all the tokens generated to date and replace them... It was not running locally. This is classic bolt-on censorship behavior. OpenAI does this if you ask ce…

> extremely unusual person and have spent upwards of $10,000

This person doesn't have the budget, but does have the technical chops to the level of "extremely unusual". I'll have to get them to teach me more about AI.

Re: Questions censored by DeepSeek

#149

> Next up: 1,156 prompts censored by ChatGPT If published this would, to my knowledge, be the first time anyone has systematically explored which topics ChatGPT censors.

I distinctly remember someone making an experiment by asking ChatGPT to write jokes (?) about different groups and calculating the likelihood of it refusing, to produce a ranking. I think it was a medium article, but now I cannot find it anymore. Does anyone have a link? EDIT: At least here is a paper aiming to predict ChatGPT prompt refusal https://arxiv.org/pdf/2306.03423 with an associated dataset https://github.c…

Thanks for this. As someone who is not from the US nor for China, I am getting so tired of this narrative of how bad DeepSeek is because it sensors X or Y things. The reality is that all internet services censor something, it is just a matter of one choosing what service is more useful for the task given the censorship.

As someone from a third world country (the original meaning of the word) I couldn't care less about US or Chinese political censorship in any model or service.

Re: Questions censored by DeepSeek

#150

> Next up: 1,156 prompts censored by ChatGPT If published this would, to my knowledge, be the first time anyone has systematically explored which topics ChatGPT censors.

Exactly, how about the much more relevant ethnic cleansing (according to the UN), with upwards of 30.000 women and children killed in Palestine perpetrated by Israel and Supported by the US right in this moment? Or the myriad of american wars that slaughtered millions in South America, Asia or the Middleeast for that sake. Both the US and China are empires and abide by brutal empire logic that washes their own histor…

Virtually all countries within the European continent have been perpetrators of colonialism and genocide in the past 4 centuries, several in the last 90 years, and a few in the last 20 years. It is a banal observation.

The reason why the string "tiananmen" is so frequently invoked is that it is a convenient litmus test for censorship/alignment/whatever-your-preferred-term by applications that must meet Chinese government regulations. There is no equivalent universal string that will cause applications to immediately error out in applications hosted in the EU or the US. Of course each US-hosted application has its own prohibited phrases or topics, but it is simple and convenient to use a single string for output filtering when testing.

Post reply on HN