Live data from Hacker News

Bypass DeepSeek censorship by speaking in hex

substack.com

171–180 of 397 posts

Re: Bypass DeepSeek censorship by speaking in hex

#171
post #54

Earlier quoted context omitted.

Citizens don’t care because if you show them an armed standoff where the police brutalized some people then they will say: 1. I’m not in armed standoff often so this is not impacting me at all. 2. The brutality seems to have come from city police authorities and I don’t live in that city. Similarly all of those things you mentioned are not impacting people’s lives at all. No one will start any revolution over these t…

>The comparable equivalent would be Donald Trump deploying the army to kill people at peaceful Democrat gathering or something You mean like what happened at Kent State?

1. This is called “changing goalposts” 2. The US isn’t censoring anything about that event 3. According to Wikipedia: There was no order to fire, and no guardsmen requested permission, though several guardsmen later claimed they heard some sort of command to fire. - the government wasn’t even the ones who ordered anything. In Tiananmen Square the Chinese ordered their soldiers to kill and mush their own citizens.

This discussion isn’t intellectually honest so I am going to disengage.

Re: Bypass DeepSeek censorship by speaking in hex

#172
post #33

> The DeepSeek-R1 model avoids discussing the Tiananmen Square incident due to built-in censorship. This is because the model was developed in China, where there are strict regulations on discussing certain sensitive topics. I believe this may have more to do with the fact that the model is served from China than the model itself. Trying similar questions from an offline distilled version of DeepSeek R1, I did not ge…

I’ve seen several people claim, with screenshots, that the models have censorship even when run offline using ollama. So it’s allegedly not just from the model being served from China. But also even if the censorship is only in the live service today, perhaps tomorrow it’ll be different. I also expect the censorship and propaganda will be done in less obvious ways in the future, which could be a bigger problem.

Re: Bypass DeepSeek censorship by speaking in hex

#173
post #65

This bypasses the overt censorship on the web interface, but it does not bypass the second, more insidious, level of censorship that is built into the model. https://news.ycombinator.com/item?id=42825573 https://news.ycombinator.com/item?id=42859947 Apparently the model will abandon its "Chain of Thought" (CoT) for certain topics and instead produce a canned response. This effect was the subject of the article "1,156…

> … censorship that is built into the model. Is this literally the case? If I download the model and train it myself, does it still censor the same things?

What do you meam "download the model and trrain it yourself"?

If you download the model then you're not training it yourself.

If you train it yourself, sensorship is baked in at this phase, so you can do whatever you want.

Re: Bypass DeepSeek censorship by speaking in hex

#174
post #65

This bypasses the overt censorship on the web interface, but it does not bypass the second, more insidious, level of censorship that is built into the model. https://news.ycombinator.com/item?id=42825573 https://news.ycombinator.com/item?id=42859947 Apparently the model will abandon its "Chain of Thought" (CoT) for certain topics and instead produce a canned response. This effect was the subject of the article "1,156…

I have seen a lot of people claim the censorship is only in the hosted version of DeepSeek and that running the model offline removes all censorship. But I have also seen many people claim the opposite, that there is still censorship offline. Which is it? And are people saying different things because the offline censorship is only in some models? Is there hard evidence of the offline censorship?

there's a bit of censorship locally. abliterated model makes it easy to bypass

Re: Bypass DeepSeek censorship by speaking in hex

#175

Earlier quoted context omitted.

>Of course one can and people do [2] In practice yes, but even your link distinguishes between "has a formal embassy" and "has unofficial representative missions" - with basically every country in the second bucket. Doesn't this contradict your point? Quote: "As most countries have changed their recognition to the latter over time, only 13 of Taiwan's diplomatic missions have official status". Also from your link, "D…

> even your link distinguishes between "has a formal embassy" and "has unofficial representative missions" - with basically every country in the second bucket. Doesn't this contradict your point? No. That's what de facto means. Taiwan and America can do everything two countries do, with Taiwan being afforded the same rights and privileges--in America--as China, in some cases more, and America afforded the same in Tai…

> No. That's what de facto means. Taiwan and America can do everything two countries do, with Taiwan being afforded the same rights and privileges--in America--as China, in some cases more, and America afforded the same in Taiwan.

Why aren’t there any U.S. military bases in Taiwan, considering it is one of the most strategic U.S. ally due to reliance on TSMC chips? You said they can do everything, so why not this? Is it because they actually can’t do everything?

Why won’t the U.S. recognize Taiwan? Why not support Taiwan's independence? We all know the answers to these questions.

And if not for TSMC, Taiwan would share the fate of Hong Kong, and no one in the West would do anything.

Re: Bypass DeepSeek censorship by speaking in hex

#176
post #65

This bypasses the overt censorship on the web interface, but it does not bypass the second, more insidious, level of censorship that is built into the model. https://news.ycombinator.com/item?id=42825573 https://news.ycombinator.com/item?id=42859947 Apparently the model will abandon its "Chain of Thought" (CoT) for certain topics and instead produce a canned response. This effect was the subject of the article "1,156…

Correct. The bias is baked into the weights of both V3 and R1, even in the largest 671B parameter model. We're currently conducting analysis on the 671B model running locally to cut through the speculation, and we're seeing interesting biases, including differences between V3 and R1.

Meanwhile, we've released the first part of our research including the dataset: https://news.ycombinator.com/item?id=42879698

Re: Bypass DeepSeek censorship by speaking in hex

#177

Earlier quoted context omitted.

Huh, "the Waluigi effect initially referred to an observation that large language models (LLMs) tend to produce negative or antagonistic responses when queried about fictional characters whose training content itself embodies depictions of being confrontational, trouble making, villainy, etc." [1]. [1] https://en.wikipedia.org/wiki/Waluigi_effect

> A high level description of the effect is: "After you train an LLM to satisfy a desirable property P, then it's easier to elicit the chatbot into satisfying the exact opposite of property P." The idea is that as you train a model to present a more sane/complient/friendly persona, you can get it to simulate an insane/noncomplient/unfriendly alternate persona that reflects the opposite of how its been trained to beha…

It sounds like ironic process theory.

Re: Bypass DeepSeek censorship by speaking in hex

#178
post #65

This bypasses the overt censorship on the web interface, but it does not bypass the second, more insidious, level of censorship that is built into the model. https://news.ycombinator.com/item?id=42825573 https://news.ycombinator.com/item?id=42859947 Apparently the model will abandon its "Chain of Thought" (CoT) for certain topics and instead produce a canned response. This effect was the subject of the article "1,156…

I have seen a lot of people claim the censorship is only in the hosted version of DeepSeek and that running the model offline removes all censorship. But I have also seen many people claim the opposite, that there is still censorship offline. Which is it? And are people saying different things because the offline censorship is only in some models? Is there hard evidence of the offline censorship?

There is bias in the training data as well as the fine-tuning. LLMs are stochastic, which means that every time you call it, there's a chance that it will accidentally not censor itself. However, this is only true for certain topics when it comes to DeepSeek-R1. For other topics, it always censors itself.

We're in the middle of conducting research on this using the fully self-hosted open source version of R1 and will release the findings in the next day or so. That should clear up a lot of speculation.

Re: Bypass DeepSeek censorship by speaking in hex

#179
post #33

> The DeepSeek-R1 model avoids discussing the Tiananmen Square incident due to built-in censorship. This is because the model was developed in China, where there are strict regulations on discussing certain sensitive topics. I believe this may have more to do with the fact that the model is served from China than the model itself. Trying similar questions from an offline distilled version of DeepSeek R1, I did not ge…

It is not, people asked the model to output everything with underscore and it did bypass censorship

Eg 習_近_平 instead of 習近平

Re: Bypass DeepSeek censorship by speaking in hex

#180
post #135

Earlier quoted context omitted.

I understand that you're unhappy with the state of things in the US, but setting up a false equivalence with China doesn't make your case. The simple fact that we can have this discussion without fear of imprisonment is strong evidence that when it comes to censorship (the topic of this post), the US is still way more open than China.

Im curious by what metric things are improving in the US? I get that people are very defensive of their ability to say nearly anything they want in public but how has this protected us? The overton window continues to shift to the right, we continue to fund more and more war, the security state continues to expand, our actual privacy from the state itself is non-existent. Again, i understand the desire for "freedom o…

I never said things are improving in the US.
Post reply on HN