Qwen 3.8 Distilled From GPT-5.5 Pro? The Prefill Evidence
The same researchers behind the stolen-thoughts decryption method ran a new experiment: seed a model's reasoning channel with the first 1% of GPT-5.5 Pro's recovered chain of thought, let it finish the answer freely, and measure how much of the teacher's answer shows up in the first 100 tokens. On 45 problems (15 STEM, 15 non-STEM, 15 synthetic puzzles held away from public web crawls), published overnight as v1.1 of the reasoning-prefill experiment, the deltas are not subtle.
What the numbers say
Qwen3.8 A95B moved +18.18pp toward GPT-5.5 Pro's answer when primed with its reasoning (16.79% → 34.97% source recall). On synthetic puzzles — questions that exist nowhere on the public internet — it moved +14.75pp. Kimi K3, the closest GPT-overlap baseline, moved only +4.54pp. DeepSeek V4 Flash moved −1.17pp, i.e. not at all. An earlier run showed Qwen barely budging toward Opus 4.8 under the same treatment.
The pattern: Qwen 3.8 knows how GPT-5.5 Pro thinks in a way Kimi and DeepSeek demonstrably don't. The author's read — and it's the simplest one consistent with the data — is that Qwen trained on GPT-5.5 Pro output or a close GPT relative's. That's the mechanism Reuters was already circling yesterday with unauthorized agent activity on 10+ more sites: model outputs are training corpus now, whether the teacher consents or not.
The caveat nobody is running with
The top HN objection is real and nobody has folded it into the headline number: the recovered GPT-5.5 Pro thoughts only became public on August 10, when the stolen-thoughts paper dropped. Qwen 3.8's 0902 snapshot trained after that date — so it may have simply read those specific traces on the open web, and the prefill effect could be familiarity, not distillation. The counter-argument (per the paper's Appendix B): Qwen didn't move toward Opus 4.8 traces under the same conditions, and the synthetic-puzzle generalization is hard to explain from reading alone.
My verdict: this is the first forensic instrument for distillation provenance — but it needs a control the author hasn't run yet. Replay the experiment on a model checkpoint that provably predates August 10. If the +18pp survives, Qwen has some explaining to do. If it collapses, the signal is contamination and the whole class of "overlap forensics" needs contamination controls baked in from day one. Watch for the v1.2.