Jungseob Lee Publications

Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models

arXiv preprint

Jungseob Lee1, Seungyoon Lee1, Suhyune Son1, Dongyub Jude Lee2, Sungbin Han1, Sugyeong Eo3*, Heuiseok Lim1*

1Korea University 2Zoom Communications 3Yonsei University

*Corresponding authors

Chains of thought written to reach a shown gold answer pass the correctness filter, yet fine-tuning a strong reasoning model on them sharply lowers its verifiable-reasoning accuracy.

Diagram with two parallel rows that share six stages, namely shared generator, prompt, generated chain, final-answer filter, same SFT, and student. In the top row, the standard answer-blind path, a shared LLM G receives a question with the answer hidden and generates a chain of steps that ends at the answer. In the bottom row, answer-visible rationalization, the same model receives the question together with the gold answer and generates a chain that begins at the answer. Both rows keep the chains whose final answer is correct and run SFT on the filtered chains with the same initialization and recipe, which gives a blind student in the top row and a rationalized student in the bottom row. A note between the rows says that only answer visibility differs.
The one-bit experiment. A shared generator produces chains for the same problems under an answer-blind path and an answer-visible rationalization path. Both arms use the same final-answer filter and the same SFT recipe, so the arms differ only in whether generation is answer-conditioned, the gold answer shown together with a request to reach it.

Abstract

A standard recipe for distilling the reasoning ability of large language models (LLMs) is to sample chains of thought from the model, keep those that reach the correct final answer, and fine-tune on the survivors. When sampling fails, a common fix shows the generator the gold answer and asks it to write a chain that reaches that answer. We show that this second step degrades the training data in a way that correctness filtering cannot catch. We run a controlled experiment that fixes the generator, the problem set, and the correctness filter, and varies only whether the chain is generated under answer-conditioning, the gold answer shown with a request to reach it. Training a strong instruction-tuned reasoning model on its own answer-conditioned chains sharply lowers its verifiable-reasoning accuracy. The loss grows with difficulty, reaching as much as about 27 points on the hardest competition problems. The mechanism is legible in the chains themselves, which rationalize backward from the shown answer instead of deriving it, with the early final-answer statement as the measurable symptom. The harm is a property of the data rather than the generator, read off unlabeled generations before any fine-tuning, ordering the penalty across eight thinking models from four families, and transferring across teacher families. A prompt ablation localizes it to the rationalize-toward instruction rather than the answer's bare visibility. The practical takeaway is to generate answer-blind, because no correctness filter can see this damage in the data. Code is available at https://github.com/js-lee-AI/answer-leakage.

The One-Bit Experiment

A single model generates chains of thought for the same problems under two prompts that differ only in whether generation is answer-conditioned.

  • Blind arm

    The answer is hidden

    The Blind prompt shows only the problem and asks the model to derive the answer. Chains whose final boxed answer is correct train the answer-blind student \(S_{\mathrm{blind}}\).

  • Leaked arm

    The gold answer is shown

    The Leaked prompt appends the gold answer \(a(q)\) and asks for a reasoning trace that arrives at it. The same filter keeps its correct chains, which train the answer-leaked student \(S_{\mathrm{leaked}}\).

\[\Delta = \operatorname{acc}(S_{\mathrm{blind}}) - \operatorname{acc}(S_{\mathrm{leaked}})\]

The leakage penalty \(\Delta\) is the accuracy lost by generating the chain under answer-conditioning. The comparison is built on a fixed set of 935 math problems. Unless noted otherwise, every trained model is a full-parameter supervised fine-tune of Qwen3-8B, evaluated on MATH-500 with a 32,768-token budget, where the base model scores 95.2.

The Leakage Penalty

Generating under answer-conditioning costs 16.2 points on MATH-500 as a three-seed mean. The Blind student sits at the base model within seed noise (95.1 against 95.2), so essentially the whole gap is leakage damage. The penalty holds across four independent math benchmarks, and within MATH-500 it grows with problem difficulty.

BenchmarkBlind ↑Leaked ↑Penalty ↓
GSM8K (grade school)95.490.54.9
MATH-500 (competition)95.178.916.2
Level 195.395.30.0
Level 298.184.114.1
Level 398.483.514.9
Level 495.674.221.4
Level 589.870.918.9
Minerva (graduate STEM)48.738.510.2
AIME 2024–2025 (olympiad)71.744.427.2

Scroll sideways to see every column.

Answer-leakage penalty on Qwen3-8B, MATH-500 accuracy unless a benchmark is named, with Penalty the Blind minus Leaked gap \(\Delta\). The tinted Blind column is the recommended answer-blind arm. The rows are panel (C) of the paper's table, which spans four difficulty-ordered math benchmarks. Panels (A) and (B) follow below.

A length-matched arm, which truncates each blind chain to its matched leaked chain's length, scores 95.0, so the penalty is carried by the answer-conditioned content, not the length.

Leaked Chains State the Answer First

A chain is answer-first if the gold final answer is stated within the first 20% of the think block. With \(\mathrm{pos}(c) \in [0, 1]\) the relative position of the first gold-answer mention inside the reasoning of a chain \(c\), the answer-first rate of a corpus \(C\) is

\[\mathrm{AFR}(C) = \frac{1}{|C|} \sum_{c \in C} \mathbf{1}\bigl[\mathrm{pos}(c) < 0.2\bigr]\]

A high AFR is the measurable symptom of rationalizing toward a known answer rather than deriving it. On the matched pairs the leaked chains show a 20.8-point rise in AFR at comparable answer coverage, and the habit reappears in each student's own MATH-500 generations, where no answer is shown.

Three panels. (a) Density of chains against the first-answer position as a fraction of the think block. The leaked curve is highest near the start and lies above the blind curve over most of the shaded answer-first zone below 0.2, and the blind curve is higher from there onward. (b) Percent for the blind and leaked arms at three points, training-data AFR of 26.6 against 47.4 with a gap of 20.8, student-chain AFR of 36.5 against 48.7 with a gap of 12.2, and MATH-500 accuracy of 95.2 against 79.6 with a gap of 15.6. (c) Chain length in thousands of characters. Blind chains average about 26 in training and about 14 at test, while leaked chains stay near 20 in both.
Answer-position and length behavior of the answer-blind and answer-leaked arms in training chains and student outputs. (a) Distribution of the first gold-answer position within the think block for the answer-blind chains, generated with the answer never shown, and the answer-leaked chains, generated with the gold answer shown. (b) Answer-first rate (AFR) in the training corpus and in the trained students' own answer-blind generations, for the blind and leaked arms. (c) Mean think-block length of the blind and leaked arms in the training chains and in the students' own test generations.

What Carries the Harm

Panel (A) holds the answer visible across all arms and varies only the instruction. The penalty falls with the answer-first rate to under a point, so the harm tracks the instruction, not the answer's mere visibility. The derive-first instruction recovers about two-thirds of the penalty.

SettingBlind ↑Leaked ↑Penalty ↓
(A) Instruction ablation at fixed visibility
Rationalize toward answer95.178.916.2
Derive-first95.190.74.3
Final-check95.193.71.3
Out-of-band95.194.30.7
(B) Within-corpus chain-population carrier
Leaked-Early94.064.229.8
Leaked-Late94.886.28.6
Difference-in-differences+16.9

Scroll sideways to see every column.

Panels (A) and (B) of the same table, MATH-500 accuracy on Qwen3-8B. (A) varies the generation instruction at fixed answer visibility. (B) is the within-corpus two-by-two carrier design.

Panel (B) splits the 935 leaked chains by whether the answer is stated within the first 20% of the think block, giving Leaked-Early (438 chains) and Leaked-Late (497), and fine-tunes separately on each. The model trained on rationalized leaked chains collapses while the derivation-first arm largely holds. Two control arms on the blind chains of the same problems turn the comparison into a difference-in-differences that nets out problem-subset effects, with a three-seed mean of +16.9 (seed range 12 to 21). The rationalized chains therefore carry most but not all of the harm, and a chain-level filter is a cheap partial repair that leaves a residue only blind generation closes.

Anticipating the Penalty Across Models

Whether answer conditioning harms a model is anticipated, before any fine-tuning, by \(\Delta\mathrm{AFR} = \mathrm{AFR}_{\text{leaked}} - \mathrm{AFR}_{\text{blind}}\), which measures how much more often the model states the gold answer early when it can see it. Computing it requires only a small unlabeled sample of generations from the candidate teacher.

Scatter plot of leakage penalty against the rationalization signature, delta AFR in percentage points, for eight models, with a dashed linear fit and a shaded band. The in-sample points are DS-Llama-8B near zero on both axes, DS-Qwen-7B and GLM-Z1-9B near a signature of 11 with penalties of about 7 and 12, Qwen3-8B near 21 with a penalty of about 16, and Qwen3-4B near 27 with a penalty of about 29. The held-out points are DS-Qwen-14B near 9 with a penalty of about 9, Llama-Nemotron-8B near 18 with a penalty of about 18, and Qwen3-1.7B near 26 with a penalty of about 30. The plot is annotated r = 0.96, n = 8.
The ΔAFR signature against the answer-leakage penalty across eight thinking models from four families. The dashed line is the linear fit and the shaded band its dispersion. Three points are out-of-sample, and labels are colored by in-sample against held-out. Error bars show sampling variability on both axes.

Across all eight models the signature explains the penalty closely at \(r = 0.960\). An independent third family, Llama-Nemotron-8B, is predicted to lose 17.4 points and loses 17.6. With only eight models, the fit should be read as a strong but small-sample trend rather than a precisely estimated law.

ModelPre-FT signaturePost-FT outcome
Base ↑ΔAFRBlind ↑Leaked ↑Penalty ↓
Held-in models
Qwen3-8B95.2+20.895.178.916.2
Qwen3-4B96.0+26.694.064.829.2
GLM-Z1-9B94.4+11.194.882.612.2
DS-Qwen-7B79.6+11.089.782.47.3
DS-Llama-8B77.8−0.265.864.61.2
Held-out models
Qwen3-1.7B92.6+26.289.659.829.8
DS-Qwen-14B88.2+9.392.283.48.8
Nemotron-8B95.4+17.992.274.617.6

Scroll sideways to see every column.

The answer-first signature across eight thinking models from four families, each self-distilled blind against leaked on MATH-500. Base is un-fine-tuned accuracy, ΔAFR the pre-fine-tuning signature. The Qwen3-8B and DS-Qwen-7B rows are three-seed means, the other six single training runs.

Generalization

The penalty holds with the same sign across every generative setting and vanishes only where the task needs no multi-step derivation, the boundary the mechanism predicts.

SettingBlind ↑Leaked ↑Penalty ↓
(A) Code domain, Qwen3-8B student
MBPP+ (pass rate)68.358.89.5
HumanEval+ (pass rate)72.863.69.2
(B) Smaller student
Cross-student (1.7B from 8B)85.778.96.9
(C) Cross-family teacher, Qwen3-8B student
Nemotron94.079.214.8
GLM93.082.810.2

Scroll sideways to see every column.

The one-bit intervention carried to the code domain, a smaller student, and two further teacher families with a Qwen3-8B student. Columns are Blind and Leaked accuracy and the red Penalty is Blind − Leaked, MATH-500 accuracy for the math rows and pass rate for the code rows.

In code, where the “answer” is the reference solution and “correct” means passing the tests, answer-conditioned chains that pass tests still harm the student. On the multiple-choice benchmarks MMLU and GPQA-Diamond, the blind and leaked students sit together near the base model, the differences within noise. The single-seed rows of the table support a consistent sign rather than an exact magnitude, and the small-student transfer is the least stable setting tested.

BibTeX

@misc{lee2026answerconditioned,
  title = {Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models},
  author = {Jungseob Lee and Seungyoon Lee and Suhyune Son and Dongyub Jude Lee and Sungbin Han and Sugyeong Eo and Heuiseok Lim},
  year = {2026},
  journal = {arXiv preprint},
  eprint = {2607.14552},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2607.14552},
}