Qwen 3.8 follows GPT-5.5 Pro reasoning prefills
gist.github.com
Qwen 3.8 follows GPT-5.5 Pro reasoning prefills
1–10 of 85 posts
Re: Qwen 3.8 follows GPT-5.5 Pro reasoning prefills
#2Qwen 3.8 0902 was trained after the release of the paper on August 10, so it should have seen those specific thoughts.
Re: Qwen 3.8 follows GPT-5.5 Pro reasoning prefills
#3If I got it correct (appending B from https://stolen-thoughts.com/paper.pdf is essential) they are the authors of the well-known exploit to recover readable CoT from OpenAI and Anthropic models. They use that to find hints of distillation, by running a benchmark with a SotA model, recovering the CoT, then taking the first 1% of the CoT and running the open-source model as if that was the start of its own CoT. In the paper they found that Kimi-K3 gets a lot closer to Claude 4.8 answers when prefilled with the start of Claude 4.8 reasoning, suggesting that Claude 4.8 was used in its post-training. This blog post is the follow-up with results that suggest that Qwen3.8 was post-trained with the help of GPT-5.5 Pro (or some similarly responding GPT model, it's unclear how many models they tested)
Re: Qwen 3.8 follows GPT-5.5 Pro reasoning prefills
#4The problem with this is obviously that the only GPT 5.5 thoughts that we have access to are from stolen thought. Qwen 3.8 0902 was trained after the release of the paper on August 10, so it should have seen those specific thoughts.
Re: Qwen 3.8 follows GPT-5.5 Pro reasoning prefills
#5How does this suggest anyting of the sorts?
Re: Qwen 3.8 follows GPT-5.5 Pro reasoning prefills
#6The problem with this is obviously that the only GPT 5.5 thoughts that we have access to are from stolen thought. Qwen 3.8 0902 was trained after the release of the paper on August 10, so it should have seen those specific thoughts.
I "independently" "invented" it for the first Anthropic reasoning models because the API required you have thoughts for each assistant message. My app lets you switch AIs within a chat, and their API used to require thinking for all messages if thinking was enabled, so I needed to get a valid thinking stub to insert.
Time has flew by for me the last 3 years, but, I'd guess it's been at least 18 months. And IMHO it wasn't very complicated to work through how to do once you were dead set on making it happen. I expect it was well-known to distillers before the paper.
Re: Qwen 3.8 follows GPT-5.5 Pro reasoning prefills
#7> Qwen barely moved toward Opus 4.8 in the earlier experiment, but moved by +20.58 points toward GPT-5.5 Pro here, including a large effect on the private synthetic puzzles. The data suggest that Qwen may have learned from GPT-5.5 Pro, or from a closely related GPT model, rather than from Opus. How does this suggest anyting of the sorts?
Re: Qwen 3.8 follows GPT-5.5 Pro reasoning prefills
#8> Qwen barely moved toward Opus 4.8 in the earlier experiment, but moved by +20.58 points toward GPT-5.5 Pro here, including a large effect on the private synthetic puzzles. The data suggest that Qwen may have learned from GPT-5.5 Pro, or from a closely related GPT model, rather than from Opus. How does this suggest anyting of the sorts?
Score go up. Probability go up. Conclusion.
Re: Qwen 3.8 follows GPT-5.5 Pro reasoning prefills
#9(slibhb: Don't get me wrong, I agree with you 99%. But the frontier labs have zero moral authority here.)
Re: Qwen 3.8 follows GPT-5.5 Pro reasoning prefills
#10News flash: people who scraped the Internet without permission to build their product complain when something vaguely similar is done to them. Water still wet, sky still blue. Film at 11. (slibhb: Don't get me wrong, I agree with you 99%. But the frontier labs have zero moral authority here.)