Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models
arXiv preprint
1Korea University 2Zoom Communications 3Yonsei University
*Corresponding authors
Abstract
A standard recipe for distilling the reasoning ability of large language models (LLMs) is to sample chains of thought from the model, keep those that reach the correct final answer, and fine-tune on the survivors. When sampling fails, a common fix shows the generator the gold answer and asks it to write a chain that reaches that answer. We show that this second step degrades the training data in a way that correctness filtering cannot catch. We run a controlled experiment that fixes the generator, the problem set, and the correctness filter, and varies only whether the chain is generated under answer-conditioning, the gold answer shown with a request to reach it. Training a strong instruction-tuned reasoning model on its own answer-conditioned chains sharply lowers its verifiable-reasoning accuracy. The loss grows with difficulty, reaching as much as about 27 points on the hardest competition problems. The mechanism is legible in the chains themselves, which rationalize backward from the shown answer instead of deriving it, with the early final-answer statement as the measurable symptom. The harm is a property of the data rather than the generator, read off unlabeled generations before any fine-tuning, ordering the penalty across eight thinking models from four families, and transferring across teacher families. A prompt ablation localizes it to the rationalize-toward instruction rather than the answer's bare visibility. The practical takeaway is to generate answer-blind, because no correctness filter can see this damage in the data. Code is available at https://
The One-Bit Experiment
A single model generates chains of thought for the same problems under two prompts that differ only in whether generation is answer-conditioned.
-
Blind arm
The answer is hidden
The Blind prompt shows only the problem and asks the model to derive the answer. Chains whose final boxed answer is correct train the answer-blind student \(S_{\mathrm{blind}}\).
-
Leaked arm
The gold answer is shown
The Leaked prompt appends the gold answer \(a(q)\) and asks for a reasoning trace that arrives at it. The same filter keeps its correct chains, which train the answer-leaked student \(S_{\mathrm{leaked}}\).
\[\Delta = \operatorname{acc}(S_{\mathrm{blind}}) - \operatorname{acc}(S_{\mathrm{leaked}})\]
The leakage penalty \(\Delta\) is the accuracy lost by generating the chain under answer-conditioning. The comparison is built on a fixed set of 935 math problems. Unless noted otherwise, every trained model is a full-parameter supervised fine-tune of Qwen3-8B, evaluated on MATH-500 with a 32,768-token budget, where the base model scores 95.2.
The Leakage Penalty
Generating under answer-conditioning costs 16.2 points on MATH-500 as a three-seed mean. The Blind student sits at the base model within seed noise (95.1 against 95.2), so essentially the whole gap is leakage damage. The penalty holds across four independent math benchmarks, and within MATH-500 it grows with problem difficulty.
| Benchmark | Blind ↑ | Leaked ↑ | Penalty ↓ |
|---|---|---|---|
| GSM8K (grade school) | 95.4 | 90.5 | 4.9 |
| MATH-500 (competition) | 95.1 | 78.9 | 16.2 |
| Level 1 | 95.3 | 95.3 | 0.0 |
| Level 2 | 98.1 | 84.1 | 14.1 |
| Level 3 | 98.4 | 83.5 | 14.9 |
| Level 4 | 95.6 | 74.2 | 21.4 |
| Level 5 | 89.8 | 70.9 | 18.9 |
| Minerva (graduate STEM) | 48.7 | 38.5 | 10.2 |
| AIME 2024–2025 (olympiad) | 71.7 | 44.4 | 27.2 |
Scroll sideways to see every column.
Answer-leakage penalty on Qwen3-8B, MATH-500 accuracy unless a benchmark is named, with Penalty the Blind minus Leaked gap \(\Delta\). The tinted Blind column is the recommended answer-blind arm. The rows are panel (C) of the paper's table, which spans four difficulty-ordered math benchmarks. Panels (A) and (B) follow below.
A length-matched arm, which truncates each blind chain to its matched leaked chain's length, scores 95.0, so the penalty is carried by the answer-conditioned content, not the length.
Leaked Chains State the Answer First
A chain is answer-first if the gold final answer is stated within the first 20% of the think block. With \(\mathrm{pos}(c) \in [0, 1]\) the relative position of the first gold-answer mention inside the reasoning of a chain \(c\), the answer-first rate of a corpus \(C\) is
\[\mathrm{AFR}(C) = \frac{1}{|C|} \sum_{c \in C} \mathbf{1}\bigl[\mathrm{pos}(c) < 0.2\bigr]\]
A high AFR is the measurable symptom of rationalizing toward a known answer rather than deriving it. On the matched pairs the leaked chains show a 20.8-point rise in AFR at comparable answer coverage, and the habit reappears in each student's own MATH-500 generations, where no answer is shown.
What Carries the Harm
Panel (A) holds the answer visible across all arms and varies only the instruction. The penalty falls with the answer-first rate to under a point, so the harm tracks the instruction, not the answer's mere visibility. The derive-first instruction recovers about two-thirds of the penalty.
| Setting | Blind ↑ | Leaked ↑ | Penalty ↓ |
|---|---|---|---|
| (A) Instruction ablation at fixed visibility | |||
| Rationalize toward answer | 95.1 | 78.9 | 16.2 |
| Derive-first | 95.1 | 90.7 | 4.3 |
| Final-check | 95.1 | 93.7 | 1.3 |
| Out-of-band | 95.1 | 94.3 | 0.7 |
| (B) Within-corpus chain-population carrier | |||
| Leaked-Early | 94.0 | 64.2 | 29.8 |
| Leaked-Late | 94.8 | 86.2 | 8.6 |
| Difference-in-differences | +16.9 | ||
Scroll sideways to see every column.
Panels (A) and (B) of the same table, MATH-500 accuracy on Qwen3-8B. (A) varies the generation instruction at fixed answer visibility. (B) is the within-corpus two-by-two carrier design.
Panel (B) splits the 935 leaked chains by whether the answer is stated within the first 20% of the think block, giving Leaked-Early (438 chains) and Leaked-Late (497), and fine-tunes separately on each. The model trained on rationalized leaked chains collapses while the derivation-first arm largely holds. Two control arms on the blind chains of the same problems turn the comparison into a difference-in-differences that nets out problem-subset effects, with a three-seed mean of +16.9 (seed range 12 to 21). The rationalized chains therefore carry most but not all of the harm, and a chain-level filter is a cheap partial repair that leaves a residue only blind generation closes.
Anticipating the Penalty Across Models
Whether answer conditioning harms a model is anticipated, before any fine-tuning, by \(\Delta\mathrm{AFR} = \mathrm{AFR}_{\text{leaked}} - \mathrm{AFR}_{\text{blind}}\), which measures how much more often the model states the gold answer early when it can see it. Computing it requires only a small unlabeled sample of generations from the candidate teacher.
Across all eight models the signature explains the penalty closely at \(r = 0.960\). An independent third family, Llama-Nemotron-8B, is predicted to lose 17.4 points and loses 17.6. With only eight models, the fit should be read as a strong but small-sample trend rather than a precisely estimated law.
| Model | Pre-FT signature | Post-FT outcome | |||
|---|---|---|---|---|---|
| Base ↑ | ΔAFR | Blind ↑ | Leaked ↑ | Penalty ↓ | |
| Held-in models | |||||
| Qwen3-8B | 95.2 | +20.8 | 95.1 | 78.9 | 16.2 |
| Qwen3-4B | 96.0 | +26.6 | 94.0 | 64.8 | 29.2 |
| GLM-Z1-9B | 94.4 | +11.1 | 94.8 | 82.6 | 12.2 |
| DS-Qwen-7B | 79.6 | +11.0 | 89.7 | 82.4 | 7.3 |
| DS-Llama-8B | 77.8 | −0.2 | 65.8 | 64.6 | 1.2 |
| Held-out models | |||||
| Qwen3-1.7B | 92.6 | +26.2 | 89.6 | 59.8 | 29.8 |
| DS-Qwen-14B | 88.2 | +9.3 | 92.2 | 83.4 | 8.8 |
| Nemotron-8B | 95.4 | +17.9 | 92.2 | 74.6 | 17.6 |
Scroll sideways to see every column.
The answer-first signature across eight thinking models from four families, each self-distilled blind against leaked on MATH-500. Base is un-fine-tuned accuracy, ΔAFR the pre-fine-tuning signature. The Qwen3-8B and DS-Qwen-7B rows are three-seed means, the other six single training runs.
Generalization
The penalty holds with the same sign across every generative setting and vanishes only where the task needs no multi-step derivation, the boundary the mechanism predicts.
| Setting | Blind ↑ | Leaked ↑ | Penalty ↓ |
|---|---|---|---|
| (A) Code domain, Qwen3-8B student | |||
| MBPP+ (pass rate) | 68.3 | 58.8 | 9.5 |
| HumanEval+ (pass rate) | 72.8 | 63.6 | 9.2 |
| (B) Smaller student | |||
| Cross-student (1.7B from 8B) | 85.7 | 78.9 | 6.9 |
| (C) Cross-family teacher, Qwen3-8B student | |||
| Nemotron | 94.0 | 79.2 | 14.8 |
| GLM | 93.0 | 82.8 | 10.2 |
Scroll sideways to see every column.
The one-bit intervention carried to the code domain, a smaller student, and two further teacher families with a Qwen3-8B student. Columns are Blind and Leaked accuracy and the red Penalty is Blind − Leaked, MATH-500 accuracy for the math rows and pass rate for the code rows.
In code, where the “answer” is the reference solution and “correct” means passing the tests, answer-conditioned chains that pass tests still harm the student. On the multiple-choice benchmarks MMLU and GPQA-Diamond, the blind and leaked students sit together near the base model, the differences within noise. The single-seed rows of the table support a consistent sign rather than an exact magnitude, and the small-student transfer is the least stable setting tested.
BibTeX
@misc{lee2026answerconditioned,
title = {Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models},
author = {Jungseob Lee and Seungyoon Lee and Suhyune Son and Dongyub Jude Lee and Sungbin Han and Sugyeong Eo and Heuiseok Lim},
year = {2026},
journal = {arXiv preprint},
eprint = {2607.14552},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2607.14552},
}