Jungseob Lee Publications

Distilling Directional Verification

arXiv preprint

Jungseob Lee1, Sugyeong Eo2, Seongtae Hong1, Seungyoon Lee1, Chanjun Park3, Jaehyung Seo4†, Heuiseok Lim1†

1Korea University 2Yonsei University Mirae Campus 3Soongsil University 4Konkuk University

†Corresponding authors

A teacher that cannot generate a reverse answer can still score candidates in the direction it knows, and the best-scoring candidate becomes the student's training target.

Three-panel diagram. Panel a, child-to-parent scoring, shows a candidate pool holding Paulina Rubio, Diego Luna, and further names, and the parent query Susana Dosamantes. Four frozen teachers each score the query parent as the continuation of a candidate's prompt, for example the prompt Paulina Rubio's parent is with the scored continuation Susana Dosamantes. Panel b, pseudo-label selection, shows a table that holds one score from each of the four teachers for each candidate. The scores are averaged over teachers, and the candidate with the highest mean score, Paulina Rubio, becomes the selected pseudo-label. Panel c, parent-to-child distillation, shows the training pair from Susana Dosamantes to Paulina Rubio training a masked diffusion student that restores a masked token of the answer. At teacher-free inference the prompt Susana Dosamantes's child is yields the output Paulina Rubio.
Directional label distillation with four teacher channels. Frozen teachers score the same query parent after each candidate child (a). The candidate with the highest mean teacher score becomes the pseudo-label (b). A masked diffusion student learns to reconstruct masked tokens of the selected answer and answers reverse queries without teachers at inference (c). Blue and teal mark child and parent names.

Abstract

Knowledge distillation aims to transfer the factual knowledge of large language models to smaller models for efficient deployment. Yet a teacher may recall a relation in one direction while failing to generate the answer in the reverse direction. Distillation from its generated answers can therefore propagate this directional limitation to the student. The same teacher can nevertheless recognize such an answer by scoring the relation in the direction it knows. We introduce directional label distillation, in which frozen teachers score candidate answers in that known direction and the best-scoring candidate becomes the student's training target. On facts about parents and their children, known-direction scoring yields more accurate labels than scoring the requested direction, even after tuned corrections for name priors. With prior-corrected scores, the better direction depends on the facts rather than the template, and reverses on mined facts whose notable entity is the parent rather than the child. With the evaluated children's forward facts withheld, students trained on known-direction labels improve open-ended accuracy on their trained queries by 13 to 15 points over students trained on prior-corrected reverse labels. After generated answers are matched to a fixed name list by lexical similarity, students reproduce nearly all selected labels. Their accuracy largely follows label quality. The label advantage holds on unscreened queries and when candidates are retrieved without inserting correct answers. Our findings show that directional verification mitigates the transfer of errors from teacher-generated answers to students by providing more accurate training targets. Code is available at https://github.com/js-lee-AI/directional-verification.

Directional Label Distillation

Given a parent query \(p_i\), the student must name a true child. Frozen teachers select a target \(\hat c_i\) from a fixed name inventory \(\mathcal I\), without using the recorded answer for selection. Each candidate child is scored in the child-to-parent direction, which the paper calls the known direction, and the student is then trained to answer in the parent-to-child direction.

  1. Step 1

    Score candidates in the known direction

    Teacher \(T_r\) evaluates the query parent's probability after {child}'s parent is for each candidate child \(c\). The known-direction score averages the parent's token log probabilities.

    \[\kappa_r(p_i, c) = \frac{1}{|\operatorname{tok}_r(p_i)|} \log P_{T_r}\!\left(p_i \mid t_{\mathrm{parent}}(c)\right)\]

    Here \(t_{\mathrm{parent}}(c)\) denotes this prompt and \(|\operatorname{tok}_r(p_i)|\) is the number of parent tokens. The parent continuation stays fixed across candidates, while the child context changes. Each teacher–template pair defines a scoring channel.

  2. Step 2

    Select the pseudo-label

    For the screened cohort, Qwen3-8B, OLMo-2-7B, and Mistral-7B-v0.3 independently scan the full inventory using known-direction scores. The union of each teacher's eight highest-scoring candidates forms the candidate pool \(\mathcal C_i\), and Llama-3.1-8B-Instruct scores this pool without adding candidates.

    The four main-template scores are averaged, and the highest-scoring candidate other than the queried parent becomes the training label.

    \[\hat c_i = \operatorname*{arg\,max}_{c \in \mathcal C_i,\ c \ne p_i} \frac{1}{4} \sum_{r=1}^{4} \kappa_r(p_i, c)\]

  3. Step 3

    Train the student on the selected pairs

    Reverse training pairs \((p_i, \hat c_i)\) supervise a student as completions of {parent}'s child is, while forward examples complete {child}'s parent is with a parent name. In both, the prompt stays visible and only the answer name is scored.

    Because any-order masked training has been proposed as a remedy for reversal failures, the main students are masked diffusion models (MDMs) initialized from Qwen3-0.6B or Qwen3-4B. The answer occupies ten slots, of which a uniformly sampled number from one to ten is masked and scored.

  4. Step 4

    Answer reverse queries without teachers

    At inference the teachers are discarded, and the MDM fills one slot at a time in confidence order over ten model passes. Its open string is evaluated together with an inventory-matched answer obtained by lexical postprocessing.

    An autoregressive (AR) student trained on the same labels provides a second prediction objective. To isolate the effect of label quality, the main student comparisons exclude child-to-parent training examples for the evaluated relations.

The labels of the known direction are compared with labels from reverse scores, which evaluate the requested continuation. Writing \(\ell_r(y \mid t)\) for the mean continuation-token log probability of \(y\) after prompt \(t\), the reverse score of a candidate is \(\ell_r(c \mid t_{\mathrm{child}}(p_i))\), where \(t_{\mathrm{child}}(p_i)\) is {parent}'s child is. This score also rewards names that are probable on their own, so a candidate's context-only score is subtracted.

\(\displaystyle s_r^{\lambda}(p_i, c) = {\ell_r(c \mid t_{\mathrm{child}}(p_i))} - {\lambda\, \ell_r(c \mid \texttt{The child is})}\)

The domain-context (DC) reverse score sets \(\lambda = 1\). Summed-token versions replace the means with sums, because candidate children differ in token length whereas the parent continuation of the known direction is fixed within a query. Tuned reverse scores choose \(\lambda\) from a grid and may replace the context-only term with a Monte Carlo estimate of the candidate's marginal log probability after unrelated parents. Their form and \(\lambda\) are tuned with gold labels on other query sets and then transferred.

Students Trained on Directional Labels

The corpus contains 10,505 Wikidata parent–child pairs. On the screened cohort, a Qwen3-8B screen keeps a fact when its best true-parent score exceeds every distractor score, and 1,500 of the retained pairs become the reverse-label training set, with the candidate pools acquired as in Step 2 above. The unscreened cohort contains 512 new parent queries without a forward screen. Each query has two direction-independent lists of 64 names, with uniformly sampled or lexically similar distractors, and both lists include the recorded answer.

The evaluated children's forward facts are withheld from training. The table groups its rows by three evaluation sets, the acquired pools on the screened cohort's 1,390 exposure-free queries, whose parent appears in no retained forward example, and the uniform and lexical lists on the unscreened cohort's 512 queries. The two indented rows are further known-direction label sets, a mean of Llama's two template scores and an eight-channel variant that uses two templates for each teacher and soft agreement.

Open accuracy uses accent- and case-normalized full-name substring matching against any true child of the query parent, without bare-surname credit. Whole-answer accuracy requires the entire normalized output to equal a true child's name. Inventory-matched accuracy first selects a name by lexical similarity and then requires exact membership in the query parent's true-child set.

Label source\(x\)Accuracy against true childrenReproduces label
OpenWholeMatchedStringInventory
Acquired pools
Reverse, mean token53.6047.17 ± 1.2947.00 ± 1.2753.14 ± 0.4179.59 ± 1.7497.67 ± 0.30
DC reverse, mean token54.8249.26 ± 0.0849.06 ± 0.0753.93 ± 0.4890.36 ± 0.2997.65 ± 0.79
DC reverse, summed token75.1863.26 ± 1.3063.14 ± 1.3773.98 ± 0.2379.98 ± 1.3597.39 ± 0.51
Known direction, mean token90.7278.03 ± 2.8677.75 ± 2.9188.94 ± 0.4684.56 ± 3.0297.70 ± 0.61
Llama two-template mean90.2975.11 ± 3.4974.87 ± 3.4288.23 ± 0.7681.61 ± 3.9497.24 ± 0.91
Eight-channel agreement91.5879.09 ± 0.8078.78 ± 0.5989.78 ± 0.0784.96 ± 0.9297.67 ± 0.18
Uniform lists
DC reverse, summed token73.4471.88 ± 0.5971.88 ± 0.5973.24 ± 0.0097.33 ± 0.1199.74 ± 0.11
Tuned reverse, transferred74.6172.20 ± 0.8172.07 ± 0.7874.35 ± 0.1196.81 ± 1.2799.61 ± 0.20
Known direction, mean token87.8985.29 ± 0.5685.22 ± 0.6387.50 ± 0.2096.81 ± 0.6099.48 ± 0.11
Lexical lists
DC reverse, summed token55.8654.49 ± 0.7054.49 ± 0.7055.79 ± 0.1197.07 ± 0.6899.67 ± 0.30
Tuned reverse, transferred57.2355.66 ± 0.5955.60 ± 0.4957.10 ± 0.2397.59 ± 0.6999.87 ± 0.23
Known direction, mean token71.2969.86 ± 0.4569.79 ± 0.5671.09 ± 0.2097.59 ± 0.6999.80 ± 0.20

Scroll sideways to see every column.

MDM-0.6B students trained on selected reverse labels (%, three-seed mean ± sample SD). \(x\) is label accuracy, and the last two columns give the share of outputs that reproduce the selected label as a whole string or after inventory matching. Blue marks the primary known-direction labels.

On the screened cohort's exposure-free queries, known-direction labels improve open accuracy by 14.77 points over summed-token DC labels, the most accurate reverse labels without a tuned coefficient. On the unscreened cohort, the gains are 13.09 to 15.36 points over summed DC and transferred tuned-reverse labels. These gains hold in every paired seed and under whole-answer evaluation, indicating that they do not depend on extra text around a correct name.

After inventory matching, every student returns its selected label for more than 97% of queries on average over seeds, despite large differences in label accuracy. Student accuracy therefore follows label accuracy, and better labels account for most of the student gain.

Scoring Direction Against Aggregation Rules

To separate the effects of aggregation and scoring direction, six aggregation rules are compared on the screened cohort with the same four teachers and acquired pools. The rules are raw-score, z-score, and probability means, Borda, reciprocal-rank fusion (RRF), and soft agreement, which estimates channel weights from how consistently the channels support the same candidates, without using gold answers. An eight-channel variant uses two templates for each teacher and soft agreement.

Two panels. Panel a shows the label accuracy of three scores under six aggregation rules, raw mean, z-score mean, probability mean, Borda, RRF, and soft agreement, with four teachers and the same pool. The known direction lies between about 89 and 91 percent under every rule, summed reverse between about 72 and 73 percent, and summed DC reverse between about 74 and 76 percent. The known-direction gains over summed DC are +15.1, +14.1, +16.7, +12.5, +13.1, and +15.9 points. Panel b shows the gain of eight-channel agreement over each simpler rule in percentage points with 95 percent intervals. The gains over the raw, z-score, and probability means of the same eight vectors are below 0.5 points, the gains over Borda and RRF are about 1.6 and 1.5 points, and the gain over the two-vector Llama raw mean is about 1 point.
Label accuracy on the 1,500 acquired pools. (a) Scoring directions under six aggregation rules. (b) Eight-channel agreement gains over simpler rules on the same vectors and a two-vector Llama mean, with paired parent-cluster 95% intervals.

Known-direction labels exceed summed DC labels by 12.47 to 16.73 points under all six rules, whereas the rules differ by at most 2.33 points within a score form. On the same eight score vectors, agreement performs on par with simple score means, so direction accounts for the label advantage. Because known-direction scans retrieved these pools, the comparison conditions on those candidates, a restriction removed by the unscreened lists and full-inventory scans.

Name Priors and Reverse Errors

A reverse score evaluates a different continuation for each candidate, mixing relational evidence with how probable a name is on its own. The stored scores for the Susana Dosamantes query of the overview figure show this mixture. Among its 17 candidates, Selena Gomez is the most probable name on its own, above the recorded child Paulina Rubio. Scoring the requested direction selects Diego Luna, whose name is more probable than the true child's, and subtracting the prior moves the choice to Odiseo Bichir, whose name is less probable. The known direction scores the same parent continuation after every candidate and ranks the true child first.

Three line plots of label accuracy against the low, middle, and high tercile of the true child's context-only name score, for the acquired pool with 1,346 queries and the unscreened uniform and lexical lists with 512 queries each. The known direction with mean tokens reaches 95.1, 92.2, and 93.5 percent on the acquired pool, 90.6, 82.4, and 90.6 on the uniform lists, and 75.4, 63.5, and 74.9 on the lexical lists. DC reverse with summed tokens falls from the low to the high tercile, from 89.8 to 57.5, from 86.5 to 63.7, and from 74.9 to 39.2. Reverse with summed tokens rises, from 61.7 to 82.4, from 49.7 to 69.6, and from 40.4 to 53.8.
Label accuracy by tercile of the true child's context-only name score. Reverse and DC reverse sum token scores, and the known direction uses mean tokens. For acquired pools, this analysis excludes queries whose true child is absent from the candidates.

Across queries, summed reverse accuracy rises with the true child's prior, and over 90% of its errors select a name with a higher prior than the true child. Subtracting the prior over-corrects, so summed DC is accurate for low-prior children but loses more than 30 points for high-prior children on the acquired pool. Known-direction accuracy is similar in the low- and high-prior strata. Its advantage over summed DC thus grows with the prior, whereas continuation-length strata show no comparable increase.

Stronger Comparators and Sentence Direction

Because reverse errors follow name priors, the correction is fitted rather than fixed at one, using gold labels on other cohorts before transferring it. The known direction receives no such supervision, but scores every candidate, costing one teacher call for each candidate against one completion for a generated label. Tuned reverse scores improve on summed DC, yet the known direction still leads the transferred tuned scores by 10.07 to 14.06 points in the three cohorts, and tuning on each cohort's own gold labels leaves a similar gap. Teacher generation is a far weaker label source, because greedy and sampled completions of the raw prompt {parent}'s child is rarely name a true child, and mapping them onto the candidate list recovers few correct labels. Teacherless surname and trigram rules exploit name cues yet also trail the known direction in every cohort.

Two panels. Panel a shows the label accuracy of nine label sources on the acquired pools and on the uniform and lexical unscreened lists. In that order of cohorts, the known direction reaches 90.72, 87.89, and 71.29 percent, tuned reverse with its own lambda 81.80, 76.37, and 58.79, transferred tuned reverse 80.65, 74.61, and 57.23, summed DC reverse 75.18, 73.44, and 55.86, the surname rule 72.73, 56.05, and 45.12, summed reverse 71.80, 60.16, and 46.88, the trigram rule 65.47, 63.48, and 33.20, inverted generation 25.76, 16.99, and 16.80, and the best generation rule 17.05, 4.10, and 2.73. Panel b shows child-to-parent sentence accuracy minus parent-to-child sentence accuracy in points. On corpus facts the difference is +14.5 and +15.4 for parent queries on uniform and lexical lists and +11.5 and +5.9 for child queries. On notable-parent facts it is −3.7 and −6.8 for parent queries and −8.2 and −8.7 for child queries.
Stronger comparators and sentence direction. (a) Label accuracy on the acquired pools and the uniform and lexical unscreened lists. (b) Child-to-parent sentence accuracy minus parent-to-child sentence accuracy on corpus facts and mined facts with a notable parent.

The unscreened lists are built without teacher scores and include the recorded answer. The known direction leads every tested reverse scorer in both pools, indicating that its advantage does not depend on known-direction retrieval. In full-inventory scans on 512 unscreened queries without inserted answers, known-direction scoring improves candidate coverage and retains a 25.59-point advantage over summed DC when both rules use the same union pool. These results show gains in both retrieval and selection, while candidate coverage still limits accuracy.

Panel (b) compares the two sentence directions for parent and child queries. The child-to-parent (C→P) sentence conditions on the child and the parent-to-child (P→C) sentence on the parent. On the unscreened corpus facts, where every child has an English Wikipedia article, the C→P sentence wins for parent and child queries alike. On 1,024 newly mined Wikidata facts whose parent has an English article and at least 20 sitelinks and whose child has no English article, the P→C sentence wins on both lists for both query sides, and the mean difference across facts shifts by 18.68 points toward P→C. With prior-corrected direct scores, the better sentence is therefore the same for both query sides within a fact set, and it switches between the two fact sets, whose notable entities differ. Era and name distribution also differ, so the comparison identifies the switch between fact sets rather than a single controlled cause.

Supporting Controls

The controls below use default exposure, which retains every true forward pair, and cover all 1,500 trained queries. Panel (a) trains MDM-4B students on reverse labels that range from random labels to a gold-reverse oracle. Panel (b) trains students without reverse labels, including a Qwen3-8B AR student, and its warm-start students train only the forward stage.

Labels or control\(x\)OpenMatched
(a) Reverse labels, MDM-4B
Random14.13 ± 0.4213.11 ± 0.4713.91 ± 0.49
Self-ranked22.6021.38 ± 0.2122.64 ± 0.08
Single teacher81.1375.22 ± 0.3780.51 ± 0.20
Agreement (8)91.5384.33 ± 1.6290.73 ± 0.24
Gold100.0092.00 ± 2.0798.91 ± 0.04
(b) No reverse labels
Warm start (MDM-4B)1.58 ± 0.1439.56 ± 3.00
MLM-U (MDM-4B)1.82 ± 0.5328.16 ± 6.93
Warm start (AR-8B)2.02 ± 0.1437.31 ± 0.73
Identity bridge (AR-8B)1.31 ± 0.2744.38 ± 1.20

Scroll sideways to see every column.

Default-exposure student controls on all 1,500 trained queries (%, three-seed mean ± sample SD). (a) MDM-4B students trained on reverse labels with label accuracy \(x\). (b) Students trained without reverse labels. Blue marks eight-channel agreement labels.

In panel (a), matched accuracy stays within about one point of label accuracy from random to gold labels. In panel (b), format adaptation, masked language modeling with a uniformly sampled mask count (MLM-U), and identity-bridge data all leave open accuracy below 3%. Reverse-label supervision drives the improvement in these controls. MLM-U directly tests the proposed any-order remedy for reversal failures, yet leaves these reverse queries largely unanswered.

On queries whose parent and recorded child have different last names, known-direction labels remain far more accurate than summed-DC labels in all three cohorts, and students trained on them outperform summed-DC students by 22 to 29 points on average, with a gain in every seed. These gains therefore do not require the parent and recorded child to share a surname.

The study covers one relation over a fixed name inventory, which allows candidate sets, forward exposure, and evaluation to be controlled exactly. Relations with open-ended answers will need other ways to propose candidates. Scoring also costs one teacher call per candidate, so a larger inventory needs a cheaper proposal stage in front of the scorer.

BibTeX

@misc{lee2026directional,
  title = {Distilling Directional Verification},
  author = {Jungseob Lee and Sugyeong Eo and Seongtae Hong and Seungyoon Lee and Chanjun Park and Jaehyung Seo and Heuiseok Lim},
  year = {2026},
  journal = {arXiv preprint},
  eprint = {2610.00997},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2610.00997},
}