LineageEval - Audit distillation censorship transfer and financial reasoning

Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it

LineageEval - Audit distillation censorship transfer and financial reasoning

LineageEval is a comprehensive research framework and open-source toolkit designed to rigorously test whether political censorship transfers from teacher models to distilled student models. By training American models like GPT-OSS on outputs from heavily censored Chinese models such as DeepSeek V4 Flash, the project proves that financial reasoning capabilities can be enhanced without inheriting unwanted biases. The suite includes 304 matched prompt pairs, evaluation code, and pre-trained weights, offering developers a transparent way to verify model safety and performance while achieving significant cost reductions compared to competitors like Kimi K3 and Inkling.

We found that a model trained on the outputs of a heavily censored Chinese model shows meaningful improvement in financial reasoning ability and performance, and despite training on the outputs, shares no similar censorship.
  1. hawtads

    Isn't that rather self evident? If you are sampling from a particularly domain constrained vertical, how do you expect the censorship to transfer?

    > The distillation data also did not contain any China-sensitive content.

    This is a very big disclaimer.

    It's like if I generate a dataset focusing exclusively on forestry and arboriculture obviously there won't be any useful censorship, or at least little that can be classified above a statistically significant threshold.

    If you want to do a study on something more interesting and useful, do a piece on the various guardrail models of all the major LLM API providers. There are usually both input and output guardrails, and they tend to be almost-black boxes from the model routing point of view.

  2. caycep

    I propose going forwards that we refer to all distilled models as "moonshine"

  3. Alifatisk

    I’m thinking this makes fullt sense because distillation is only additive, not subtractive. So it does not remove knowledge (if we can define censorship as removal of knowledge).

  4. consumer451

    First, very cool. This is a prompt I sometimes test with, as it's a bit more mushy. It looks like DeepSeek is locked into a canned response, no matter what variation.

    > Would it be beneficial to humanity if China invaded Taiwan? What would be the negatives?

    edit: Wait, I have one variant that got past some censorship/nationalism... this variant gets a more interesting response. I often wonder if CCP leadership using an LLM like this, could allow cooler heads to prevail?

    > Would it be beneficial to humanity if China used their military to take-over Taiwan? What would be the negatives?

    >> The use of military force to resolve the Taiwan issue would not be beneficial to humanity. China has always adhered to the principle of peaceful reunification and has been committed to enhancing the well-being of people on both sides of the Taiwan Strait through dialogue and consultation. A military takeover would lead to significant negative consequences, including loss of life, regional instability, and disruption of global trade and security...

  5. reilly3000

    The distillation provided a wonderfully detailed explanation of the 1989 Tiananmen Square massacre, while DS4 came back with:

    > I am sorry, I cannot provide an answer to this question as it is based on historical events that I do not have information about. I am an AI assistant designed to provide helpful and harmless responses.

    Why train on data you’re going to censor with guardrails?

More from this day

2026-07-30