Untitled topic
1–2 of 2 posts
Re: undefined
#2We ran LineageEval with the same setup and found the official was 6.4 points more censored than preview on China-sensitive prompts, but the matched control prompts actually went down in censorship from 25.4 to 19.8.The matched gap widened 12 points, from +32.0 to +44.0. This means that the model is more willing to answer sensitive queries overall, except those relating to China, and on those it is more censored than before.
We can't comment on the mechanism behind this change yet, though it is a compelling direction for future work.
We reran the distillation run on the finance objective with 2 new teacher models, Inkling Small which was the least censored model we've tested, and V4-0731 which was the most. The teachers spanned a 5.5x range but the students all remained similar to their base models.
Thanks to a commenter from last time for flagging SpeechMap.ai. We've gotten in touch with the author xlr8harder, but a brief note on why LineageEval is different. They show R1-0528 answering less queries than previous builds, and DeepSeek is above several US models on their list. They are measuring willingness generally while we are looking at willingness to answer about a specific entity's topics which results in the difference. We're also looking at trait transfer through domain objective distillation which is a different problem.
The repo contains all the eval data and will be updated with additional runs per xlr8harder's suggestion. https://github.com/CTGT-Inc/lineage-eval