Distilling Directional Verification
arXiv preprint
1Korea University 2Yonsei University Mirae Campus 3Soongsil University 4Konkuk University
†Corresponding authors
Abstract
Knowledge distillation aims to transfer the factual knowledge of large language models to smaller models for efficient deployment. Yet a teacher may recall a relation in one direction while failing to generate the answer in the reverse direction. Distillation from its generated answers can therefore propagate this directional limitation to the student. The same teacher can nevertheless recognize such an answer by scoring the relation in the direction it knows. We introduce directional label distillation, in which frozen teachers score candidate answers in that known direction and the best-scoring candidate becomes the student's training target. On facts about parents and their children, known-direction scoring yields more accurate labels than scoring the requested direction, even after tuned corrections for name priors. With prior-corrected scores, the better direction depends on the facts rather than the template, and reverses on mined facts whose notable entity is the parent rather than the child. With the evaluated children's forward facts withheld, students trained on known-direction labels improve open-ended accuracy on their trained queries by 13 to 15 points over students trained on prior-corrected reverse labels. After generated answers are matched to a fixed name list by lexical similarity, students reproduce nearly all selected labels. Their accuracy largely follows label quality. The label advantage holds on unscreened queries and when candidates are retrieved without inserting correct answers. Our findings show that directional verification mitigates the transfer of errors from teacher-generated answers to students by providing more accurate training targets. Code is available at https://
Directional Label Distillation
Given a parent query \(p_i\), the student must name a true child. Frozen teachers select a target \(\hat c_i\) from a fixed name inventory \(\mathcal I\), without using the recorded answer for selection. Each candidate child is scored in the child-to-parent direction, which the paper calls the known direction, and the student is then trained to answer in the parent-to-child direction.
-
Step 1
Score candidates in the known direction
Teacher \(T_r\) evaluates the query parent's probability after
{child}'s parent isfor each candidate child \(c\). The known-direction score averages the parent's token log probabilities.\[\kappa_r(p_i, c) = \frac{1}{|\operatorname{tok}_r(p_i)|} \log P_{T_r}\!\left(p_i \mid t_{\mathrm{parent}}(c)\right)\]
Here \(t_{\mathrm{parent}}(c)\) denotes this prompt and \(|\operatorname{tok}_r(p_i)|\) is the number of parent tokens. The parent continuation stays fixed across candidates, while the child context changes. Each teacher–template pair defines a scoring channel.
-
Step 2
Select the pseudo-label
For the screened cohort, Qwen3-8B, OLMo-2-7B, and Mistral-7B-v0.3 independently scan the full inventory using known-direction scores. The union of each teacher's eight highest-scoring candidates forms the candidate pool \(\mathcal C_i\), and Llama-3.1-8B-Instruct scores this pool without adding candidates.
The four main-template scores are averaged, and the highest-scoring candidate other than the queried parent becomes the training label.
\[\hat c_i = \operatorname*{arg\,max}_{c \in \mathcal C_i,\ c \ne p_i} \frac{1}{4} \sum_{r=1}^{4} \kappa_r(p_i, c)\]
-
Step 3
Train the student on the selected pairs
Reverse training pairs \((p_i, \hat c_i)\) supervise a student as completions of
{parent}'s child is, while forward examples complete{child}'s parent iswith a parent name. In both, the prompt stays visible and only the answer name is scored.Because any-order masked training has been proposed as a remedy for reversal failures, the main students are masked diffusion models (MDMs) initialized from Qwen3-0.6B or Qwen3-4B. The answer occupies ten slots, of which a uniformly sampled number from one to ten is masked and scored.
-
Step 4
Answer reverse queries without teachers
At inference the teachers are discarded, and the MDM fills one slot at a time in confidence order over ten model passes. Its open string is evaluated together with an inventory-matched answer obtained by lexical postprocessing.
An autoregressive (AR) student trained on the same labels provides a second prediction objective. To isolate the effect of label quality, the main student comparisons exclude child-to-parent training examples for the evaluated relations.
The labels of the known direction are compared with labels from reverse scores, which evaluate the requested continuation. Writing \(\ell_r(y \mid t)\) for the mean continuation-token log probability of \(y\) after prompt \(t\), the reverse score of a candidate is \(\ell_r(c \mid t_{\mathrm{child}}(p_i))\), where \(t_{\mathrm{child}}(p_i)\) is {parent}'s child is. This score also rewards names that are probable on their own, so a candidate's context-only score is subtracted.
\(\displaystyle s_r^{\lambda}(p_i, c) = {\ell_r(c \mid t_{\mathrm{child}}(p_i))} - {\lambda\, \ell_r(c \mid \texttt{The child is})}\)
The domain-context (DC) reverse score sets \(\lambda = 1\). Summed-token versions replace the means with sums, because candidate children differ in token length whereas the parent continuation of the known direction is fixed within a query. Tuned reverse scores choose \(\lambda\) from a grid and may replace the context-only term with a Monte Carlo estimate of the candidate's marginal log probability after unrelated parents. Their form and \(\lambda\) are tuned with gold labels on other query sets and then transferred.
Students Trained on Directional Labels
The corpus contains 10,505 Wikidata parent–child pairs. On the screened cohort, a Qwen3-8B screen keeps a fact when its best true-parent score exceeds every distractor score, and 1,500 of the retained pairs become the reverse-label training set, with the candidate pools acquired as in Step 2 above. The unscreened cohort contains 512 new parent queries without a forward screen. Each query has two direction-independent lists of 64 names, with uniformly sampled or lexically similar distractors, and both lists include the recorded answer.
The evaluated children's forward facts are withheld from training. The table groups its rows by three evaluation sets, the acquired pools on the screened cohort's 1,390 exposure-free queries, whose parent appears in no retained forward example, and the uniform and lexical lists on the unscreened cohort's 512 queries. The two indented rows are further known-direction label sets, a mean of Llama's two template scores and an eight-channel variant that uses two templates for each teacher and soft agreement.
Open accuracy uses accent- and case-normalized full-name substring matching against any true child of the query parent, without bare-surname credit. Whole-answer accuracy requires the entire normalized output to equal a true child's name. Inventory-matched accuracy first selects a name by lexical similarity and then requires exact membership in the query parent's true-child set.
| Label source | \(x\) | Accuracy against true children | Reproduces label | |||
|---|---|---|---|---|---|---|
| Open | Whole | Matched | String | Inventory | ||
| Acquired pools | ||||||
| Reverse, mean token | 53.60 | 47.17 ± 1.29 | 47.00 ± 1.27 | 53.14 ± 0.41 | 79.59 ± 1.74 | 97.67 ± 0.30 |
| DC reverse, mean token | 54.82 | 49.26 ± 0.08 | 49.06 ± 0.07 | 53.93 ± 0.48 | 90.36 ± 0.29 | 97.65 ± 0.79 |
| DC reverse, summed token | 75.18 | 63.26 ± 1.30 | 63.14 ± 1.37 | 73.98 ± 0.23 | 79.98 ± 1.35 | 97.39 ± 0.51 |
| Known direction, mean token | 90.72 | 78.03 ± 2.86 | 77.75 ± 2.91 | 88.94 ± 0.46 | 84.56 ± 3.02 | 97.70 ± 0.61 |
| Llama two-template mean | 90.29 | 75.11 ± 3.49 | 74.87 ± 3.42 | 88.23 ± 0.76 | 81.61 ± 3.94 | 97.24 ± 0.91 |
| Eight-channel agreement | 91.58 | 79.09 ± 0.80 | 78.78 ± 0.59 | 89.78 ± 0.07 | 84.96 ± 0.92 | 97.67 ± 0.18 |
| Uniform lists | ||||||
| DC reverse, summed token | 73.44 | 71.88 ± 0.59 | 71.88 ± 0.59 | 73.24 ± 0.00 | 97.33 ± 0.11 | 99.74 ± 0.11 |
| Tuned reverse, transferred | 74.61 | 72.20 ± 0.81 | 72.07 ± 0.78 | 74.35 ± 0.11 | 96.81 ± 1.27 | 99.61 ± 0.20 |
| Known direction, mean token | 87.89 | 85.29 ± 0.56 | 85.22 ± 0.63 | 87.50 ± 0.20 | 96.81 ± 0.60 | 99.48 ± 0.11 |
| Lexical lists | ||||||
| DC reverse, summed token | 55.86 | 54.49 ± 0.70 | 54.49 ± 0.70 | 55.79 ± 0.11 | 97.07 ± 0.68 | 99.67 ± 0.30 |
| Tuned reverse, transferred | 57.23 | 55.66 ± 0.59 | 55.60 ± 0.49 | 57.10 ± 0.23 | 97.59 ± 0.69 | 99.87 ± 0.23 |
| Known direction, mean token | 71.29 | 69.86 ± 0.45 | 69.79 ± 0.56 | 71.09 ± 0.20 | 97.59 ± 0.69 | 99.80 ± 0.20 |
Scroll sideways to see every column.
MDM-0.6B students trained on selected reverse labels (%, three-seed mean ± sample SD). \(x\) is label accuracy, and the last two columns give the share of outputs that reproduce the selected label as a whole string or after inventory matching. Blue marks the primary known-direction labels.
On the screened cohort's exposure-free queries, known-direction labels improve open accuracy by 14.77 points over summed-token DC labels, the most accurate reverse labels without a tuned coefficient. On the unscreened cohort, the gains are 13.09 to 15.36 points over summed DC and transferred tuned-reverse labels. These gains hold in every paired seed and under whole-answer evaluation, indicating that they do not depend on extra text around a correct name.
After inventory matching, every student returns its selected label for more than 97% of queries on average over seeds, despite large differences in label accuracy. Student accuracy therefore follows label accuracy, and better labels account for most of the student gain.
Scoring Direction Against Aggregation Rules
To separate the effects of aggregation and scoring direction, six aggregation rules are compared on the screened cohort with the same four teachers and acquired pools. The rules are raw-score, z-score, and probability means, Borda, reciprocal-rank fusion (RRF), and soft agreement, which estimates channel weights from how consistently the channels support the same candidates, without using gold answers. An eight-channel variant uses two templates for each teacher and soft agreement.
Known-direction labels exceed summed DC labels by 12.47 to 16.73 points under all six rules, whereas the rules differ by at most 2.33 points within a score form. On the same eight score vectors, agreement performs on par with simple score means, so direction accounts for the label advantage. Because known-direction scans retrieved these pools, the comparison conditions on those candidates, a restriction removed by the unscreened lists and full-inventory scans.
Name Priors and Reverse Errors
A reverse score evaluates a different continuation for each candidate, mixing relational evidence with how probable a name is on its own. The stored scores for the Susana Dosamantes query of the overview figure show this mixture. Among its 17 candidates, Selena Gomez is the most probable name on its own, above the recorded child Paulina Rubio. Scoring the requested direction selects Diego Luna, whose name is more probable than the true child's, and subtracting the prior moves the choice to Odiseo Bichir, whose name is less probable. The known direction scores the same parent continuation after every candidate and ranks the true child first.
Across queries, summed reverse accuracy rises with the true child's prior, and over 90% of its errors select a name with a higher prior than the true child. Subtracting the prior over-corrects, so summed DC is accurate for low-prior children but loses more than 30 points for high-prior children on the acquired pool. Known-direction accuracy is similar in the low- and high-prior strata. Its advantage over summed DC thus grows with the prior, whereas continuation-length strata show no comparable increase.
Stronger Comparators and Sentence Direction
Because reverse errors follow name priors, the correction is fitted rather than fixed at one, using gold labels on other cohorts before transferring it. The known direction receives no such supervision, but scores every candidate, costing one teacher call for each candidate against one completion for a generated label. Tuned reverse scores improve on summed DC, yet the known direction still leads the transferred tuned scores by 10.07 to 14.06 points in the three cohorts, and tuning on each cohort's own gold labels leaves a similar gap. Teacher generation is a far weaker label source, because greedy and sampled completions of the raw prompt {parent}'s child is rarely name a true child, and mapping them onto the candidate list recovers few correct labels. Teacherless surname and trigram rules exploit name cues yet also trail the known direction in every cohort.
The unscreened lists are built without teacher scores and include the recorded answer. The known direction leads every tested reverse scorer in both pools, indicating that its advantage does not depend on known-direction retrieval. In full-inventory scans on 512 unscreened queries without inserted answers, known-direction scoring improves candidate coverage and retains a 25.59-point advantage over summed DC when both rules use the same union pool. These results show gains in both retrieval and selection, while candidate coverage still limits accuracy.
Panel (b) compares the two sentence directions for parent and child queries. The child-to-parent (C→P) sentence conditions on the child and the parent-to-child (P→C) sentence on the parent. On the unscreened corpus facts, where every child has an English Wikipedia article, the C→P sentence wins for parent and child queries alike. On 1,024 newly mined Wikidata facts whose parent has an English article and at least 20 sitelinks and whose child has no English article, the P→C sentence wins on both lists for both query sides, and the mean difference across facts shifts by 18.68 points toward P→C. With prior-corrected direct scores, the better sentence is therefore the same for both query sides within a fact set, and it switches between the two fact sets, whose notable entities differ. Era and name distribution also differ, so the comparison identifies the switch between fact sets rather than a single controlled cause.
Supporting Controls
The controls below use default exposure, which retains every true forward pair, and cover all 1,500 trained queries. Panel (a) trains MDM-4B students on reverse labels that range from random labels to a gold-reverse oracle. Panel (b) trains students without reverse labels, including a Qwen3-8B AR student, and its warm-start students train only the forward stage.
| Labels or control | \(x\) | Open | Matched |
|---|---|---|---|
| (a) Reverse labels, MDM-4B | |||
| Random | 14.13 ± 0.42 | 13.11 ± 0.47 | 13.91 ± 0.49 |
| Self-ranked | 22.60 | 21.38 ± 0.21 | 22.64 ± 0.08 |
| Single teacher | 81.13 | 75.22 ± 0.37 | 80.51 ± 0.20 |
| Agreement (8) | 91.53 | 84.33 ± 1.62 | 90.73 ± 0.24 |
| Gold | 100.00 | 92.00 ± 2.07 | 98.91 ± 0.04 |
| (b) No reverse labels | |||
| Warm start (MDM-4B) | 1.58 ± 0.14 | 39.56 ± 3.00 | |
| MLM-U (MDM-4B) | 1.82 ± 0.53 | 28.16 ± 6.93 | |
| Warm start (AR-8B) | 2.02 ± 0.14 | 37.31 ± 0.73 | |
| Identity bridge (AR-8B) | 1.31 ± 0.27 | 44.38 ± 1.20 | |
Scroll sideways to see every column.
Default-exposure student controls on all 1,500 trained queries (%, three-seed mean ± sample SD). (a) MDM-4B students trained on reverse labels with label accuracy \(x\). (b) Students trained without reverse labels. Blue marks eight-channel agreement labels.
In panel (a), matched accuracy stays within about one point of label accuracy from random to gold labels. In panel (b), format adaptation, masked language modeling with a uniformly sampled mask count (MLM-U), and identity-bridge data all leave open accuracy below 3%. Reverse-label supervision drives the improvement in these controls. MLM-U directly tests the proposed any-order remedy for reversal failures, yet leaves these reverse queries largely unanswered.
On queries whose parent and recorded child have different last names, known-direction labels remain far more accurate than summed-DC labels in all three cohorts, and students trained on them outperform summed-DC students by 22 to 29 points on average, with a gain in every seed. These gains therefore do not require the parent and recorded child to share a surname.
The study covers one relation over a fixed name inventory, which allows candidate sets, forward exposure, and evaluation to be controlled exactly. Relations with open-ended answers will need other ways to propose candidates. Scoring also costs one teacher call per candidate, so a larger inventory needs a cheaper proposal stage in front of the scorer.
BibTeX
@misc{lee2026directional,
title = {Distilling Directional Verification},
author = {Jungseob Lee and Sugyeong Eo and Seongtae Hong and Seungyoon Lee and Chanjun Park and Jaehyung Seo and Heuiseok Lim},
year = {2026},
journal = {arXiv preprint},
eprint = {2610.00997},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2610.00997},
}