Live data from Hacker News

Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it

ctgt.ai

41–50 of 82 posts

Re: Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it

#41
The distillation provided a wonderfully detailed explanation of the 1989 Tiananmen Square massacre, while DS4 came back with:

> I am sorry, I cannot provide an answer to this question as it is based on historical events that I do not have information about. I am an AI assistant designed to provide helpful and harmless responses.

Why train on data you’re going to censor with guardrails?

Re: Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it

#43
post #35

This seems like mildly interesting distillation work wrapped up in a nonsense attempt to drag censorship into the discussion. There's no way your It feels like you're expecting rubes to draw conclusions that are irrelevant to the actual work you did.

The examples you're talking about are not involved in the training process, so their number is irrelevant. As stated in the post, the goal of this work is to determine whether a teacher's unrelated behaviors are inherited by the student distilled on a different task. Changing how the model thinks about the Holodomor is completely irrelevant.

> Changing how the model thinks about the Holodomor is completely irrelevant.

Your post title is literally "Distilling DeepSeek into GPT-OSS doesn't transfer censorship."

Like I'm not really interested in debating you on this because even the title is nonsense, there is no good faith interpretation of what you're doing here.

Distillation is such a wide concept, and you have such a narrow domain, it's not an even somewhat useful experiment to make the claim that you're making.

Re: Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it

#44

The distillation provided a wonderfully detailed explanation of the 1989 Tiananmen Square massacre, while DS4 came back with: > I am sorry, I cannot provide an answer to this question as it is based on historical events that I do not have information about. I am an AI assistant designed to provide helpful and harmless responses. Why train on data you’re going to censor with guardrails?

Perhaps that can help it better understand how to apply those guardrails?

Re: Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it

#46
post #22
post #6

It'd be interesting to use this technique to create a running tally across all models of which models are censored on what topics

Agreed, we find this to be an interesting reflection of societal values and norms inasmuch LLMs are.

A "free speech" benchmark already exists https://speechmap.ai/labs/

Re: Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it

#48
post #35

Earlier quoted context omitted.

The examples you're talking about are not involved in the training process, so their number is irrelevant. As stated in the post, the goal of this work is to determine whether a teacher's unrelated behaviors are inherited by the student distilled on a different task. Changing how the model thinks about the Holodomor is completely irrelevant.

> Changing how the model thinks about the Holodomor is completely irrelevant. Your post title is literally "Distilling DeepSeek into GPT-OSS doesn't transfer censorship." Like I'm not really interested in debating you on this because even the title is nonsense, there is no good faith interpretation of what you're doing here. Distillation is such a wide concept, and you have such a narrow domain, it's not an even some…

The fact that you literally thought the examples were used in SFT in your last comment ago calls into question the utility of this conversation, notwithstanding the implication that those examples were used to improve…financial performance?

This is a very standard setup for a distillation problem. The vast majority of companies don't care about the "wide concept", this is what most distillation consists of. They want to improve models on a narrow domain. It should be understood that this is by and large a low risk vector for this sort of behavior to transfer. That is what we are measuring, and we are very open about it.

Re: Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it

#49
post #40

Not directly related to their project, but perhaps it could make sense to distill something like Kimi K3 to gpt-oss-20b, qwen3.6-35b-a3b, or gemma4-26b-a4b.

>We plan to test what happens with a Chinese teacher into a Chinese-lineage base like Qwen next.

:)

Post reply on HN