Jungseob Lee Publications

Risk-Controlled Selective LLM Answering by Pricing Label-Free Checks

arXiv preprint

Dongyub Jude Lee1*, Jungseob Lee2*, Chanjun Park3, Hyeonseok Moon4†, Heuiseok Lim2†

1Zoom Communications 2Korea University 3Soongsil University 4Sookmyung Women's University

*Equal contribution †Corresponding authors

PriceCheck prices label-free checks by their agreement rates and cost, composes the prices into schedules, and selects one schedule with a calibration test at a stated selective-risk target.

Two-row diagram. Top row, three boxes joined by arrows. Price: a probe of n labelled answers gives each check a completeness p1 and a leak p0. Compose: a small check feeds a large one, and the acceptance probability is P(3 of 3) plus P(1 or 2 of 3) times f(3, 2). Certify: a table of schedules is tested with p at most delta over the family size. Bottom row, one example. A candidate answer, the proposed answer 14 to a problem asking for the least possible sum of distinct positive integers whose product is 84, goes through three small re-solves, of which 2 of 3 agree. Three of three would serve and none of three would abstain. One or two of three pass to a large race, which serves at 3 agreements before 2 refusals and abstains otherwise. Served answers carry a risk at most the target.
Overview of PriceCheck. Top, the three steps: price each check on a labelled probe, compose prices into a schedule, and certify one schedule. Bottom, one candidate answer passing through a two-stage cascade. Green serves, red abstains, and amber passes the undecided band on.

Abstract

Serving an answer from a large language model requires deciding when to abstain, yet a verifier's ranking accuracy alone does not determine the error rate among served answers. We introduce PriceCheck, which builds a compact family of decision rules from label-free checks such as re-solving a problem. Each check has a price: its agreement rates on correct and incorrect answers and its cost per run. Prices fitted on a small, class-enriched labelled set compose into predictions of a schedule's coverage and cost, guiding which checks to run and when to stop. A calibration test then selects a schedule at a stated selective-risk target. In mathematics, the selected schedules serve 76.1% of answers on average and keep held-out selective risk below 1.5% on all 15 splits. Under the shared testing protocol, PriceCheck serves more answers at that target than reward models, a prompted judge, the generator's confidence and a trained correctness classifier. At matched coverage, it keeps the fewest wrong answers among these scorers. Across 118 diagnostic schedules, price-based coverage predictions have a rank correlation of 0.97 with observed coverage. These results show that choosing how checks are combined and stopped matters alongside how well a verifier ranks answers. Code is available at https://github.com/js-lee-AI/PriceCheck.

PriceCheck

A check reads a candidate answer, but not its label, and returns a binary verdict on whether it agrees with that answer. A schedule \(S\) uses one or more checks to serve an answer or abstain. Its coverage is the share of answers served, and its selective risk \(R(S)\) is the error rate among them. Given a target \(\alpha\) and a confidence budget \(\delta\), the objective is to select and certify a schedule for which \(R(S) \le \alpha\) holds with probability at least \(1 - \delta\) over the draw of calibration data, at the largest coverage obtainable.

  1. Step 1

    Price each check

    The price of a check is two measured rates and a unit cost. Completeness \(p_1\) is the per-draw rate at which the check agrees with a correct answer, and leak \(p_0\) the rate at which it agrees with a wrong one.

    \[\begin{aligned} p_1 &= \mathbb{P}(V = 1 \mid Y = 1) \\ p_0 &= \mathbb{P}(V = 1 \mid Y = 0) \end{aligned}\]

    Here \(V\) is the verdict of one draw and \(Y\) the correctness of the answer. Both rates are fitted on a small labelled pricing set.

  2. Step 2

    Compose prices into a schedule

    Prices compose into predictions of a schedule's coverage, selective risk and cost. With \(a_y\) the fitted acceptance probability for class \(y\) and \(\pi\) the correct-answer fraction of the pricing set, predicted coverage and risk are

    \[\begin{aligned} \hat{c} &= \pi\, a_1 + (1-\pi)\, a_0 \\ \hat{r} &= \frac{(1-\pi)\, a_0}{\hat{c}} \end{aligned}\]

    Cost composes in the same way, as each stage is paid only by the answers that reach it.

  3. Step 3

    Certify one schedule

    Each schedule in a family \(G\), fixed independently of the i.i.d. calibration answers, gets an exact one-sided binomial \(p\)-value against \(H_0\colon R(S) > \alpha\). It is certified when it serves \(n > 0\) calibration answers, \(e\) of them wrong, and

    \[F(e;\, n,\, \alpha) \le \delta / |G|\]

    with \(F\) the binomial cumulative distribution function, \(\delta = 0.05\) and \(|G| = 23\). The default selector takes the certified schedule with the largest calibration coverage.

Schedules over Priced Checks

The paper calls a check an action and prices seven of them. A backward probe has a reverse model write the reasoning that leads to the proposed answer, and a forward model then recovers an answer from that trace. A re-solve vote draws a fresh solution from the frozen forward model. Two generator-confidence actions read the generator's probability of Yes when asked whether the proposed answer is correct, and the mean log-probability of the stored solution.

Backward probeRe-solve voteGenerator confidence
8B4B1.7B8B1.7Bp(True)Log-prob.
Rate per draw (%)
Completeness \(p_1\)70.173.163.980.279.572.561.9
Leak \(p_0\)10.813.812.59.010.654.530.3
Cost per draw (tokens)
Generated1034103410345384480100
Weighted10345172175384100800

Scroll sideways to see every column.

Price of each action on the frozen pool: completeness \(p_1\) and leak \(p_0\) per draw, and unit cost per draw in generated and in parameter-weighted tokens. The two confidence actions re-score the stored solution, so they generate nothing.

Schedules are built from four primitives. A threshold serves when the agreement count of one backward probe reaches \(t\), and unanimity when the first \(N\) re-solve votes all agree. An adaptive race draws re-solve votes until it reaches \(k\) agreements, where it serves, or \(j\) disagreements, where it abstains. A cascade runs a cheap action first and sends its undecided band to the race. The deployed family is a fixed grid of 23 schedules that no certification run alters.

Main Results

Five generators, all fine-tuned from one 8B backbone, each attempted the 500 problems of MATH-500, and the evaluation uses 2326 of their answers with a base error of 4.26%. PriceCheck is compared with a prompted judge, a trained process reward model (PRM), Math-Shepherd, and an outcome reward model (ORM) under the shared testing protocol across fifteen splits. For a fair comparison, each rival scorer receives a family of selective rules of the same shape and comparable size, and the comparisons apply the test at its nominal level and assess realised held-out risk.

Method\(\alpha\) = 0.5%0.75%1.0%1.25%1.5%2.0%
Spl.Cov.Spl.Cov.Spl.Cov.Spl.Cov.Spl.Cov.Spl.Cov.
PriceCheck (ours)523.201257.001572.111573.851576.091578.86
Prompted verifier
Judge00516.741552.431562.321569.501578.17
Judge†00520.091561.081567.991571.911580.81
Trained verifiers
PRM0013.31516.71826.771555.901568.14
PRM†0013.90832.091562.461566.551574.20
Math-Shepherd0000000013.56517.52
Math-Shepherd†0000000027.61623.80
ORM00000013.47310.231138.63
ORM†0000000027.66934.94

Scroll sideways to see every column.

Certification under a shared protocol. At each target \(\alpha\), splits certifying of 15 and coverage over all 15 counting non-certifying splits as zero. A dagger marks 24 levels, not six. Bold is best. The last column of the paper's table, wrong answers kept at matched coverage, is given in the text below.

PriceCheck certifies on all fifteen splits at the primary 1.5% target, with the highest mean held-out coverage, 76.09%, and no test-risk exceedance. Its coverage advantage over the best rival grows from 4.18 points at the primary target to 36.91 points at 0.75%, and at the strictest 0.5% target only PriceCheck certifies any split. The judge's finer family leads at the looser 2% target.

At matched coverage, each rival's test answers are ranked and exactly as many are kept as the certified schedule kept on that split. Every rival scorer then keeps more wrong answers than PriceCheck's 44. The count is 60 for the judge, the strongest ranker, 97 for the PRM, 164 for Math-Shepherd and 195 for the ORM. Neither generator-confidence score certifies on any split at any target, and a 1.7B correctness classifier fitted to the calibration labels certifies on only three of the fifteen splits even at its widest context window.

Across Split Schemes

The fifteen splits are five source-held-out folds, each reserving one generator, and ten question-disjoint halves that reserve problems. Certified coverage is higher on the folds than on the halves. The default selector certifies an 8B re-solve vote on every split, while a cost-minimising selector instead picks a small-first cascade on three of the five folds, where its selections average 77.7% coverage at under a third of the adaptive vote's weighted cost.

Three panels. (a) Bars of certified coverage on the five source-held-out folds, by held-out generator: F0 79.0, ForwardPref 80.7, STaR 80.6, SelfDistill 79.0 and iter2 80.0 percent, all with Race (2,3) selected. (b) Bars for the ten question-disjoint halves A to J: 71.8, 73.4, 71.5, 80.9, 68.5, 81.1, 70.3, 71.7, 79.6 and 73.3 percent. Halves D and F select Race (2,3), half I selects Race (3,3), and the other seven select the unanimous 3-vote. (c) Coverage certified by 8B re-solve agreement rises from 72.2 percent at 3 re-solves per answer to 73.6 at 4, 75.7 at 8 and 76.2 at 16, against a dashed PriceCheck line at 76.1 percent.
Held-out coverage certified by PriceCheck at the 1.5% target on (a) source-held-out folds and (b) question-disjoint halves A–J, coloured by the selected configuration. Race \((k,j)\) serves at \(k\) agreements before \(j\) disagreements. (c) Coverage certified at 1.5% by 8B re-solve agreement over \(n\) re-solves per answer, against PriceCheck (dashed).

Panel (c) certifies the plain fraction of stored re-solves agreeing with an answer as a scalar score under the same protocol. With three 8B re-solves, which generate about as many tokens as the certified schedule, this score serves less at the primary target, and it reaches the schedule's coverage only at sixteen votes.

Prices Predict Schedule Coverage

Prices fitted on 100 class-enriched answers track cascade coverage across three pools. Over 118 diagnostic schedules, price-based coverage predictions have a rank correlation of 0.97 with observed coverage, and the mean absolute coverage error is 2.97 points over 70 cascades. Risk prediction has a mean error of about one point but is less reliable across schedules, so PriceCheck tests schedules on calibration data rather than using predicted risk as a certificate.

Three panels. (a) Predicted minus realised coverage in points for the probe-first and small-first cascades: about minus 0.3 and minus 1.8 on MATH, plus 1.8 and plus 1.0 on GSM8K, and minus 3.6 and minus 5.7 on Minerva. (b) Splits over target, of 15, by risk target. At 0.5 percent the uncorrected rule exceeds on 2 splits and the plug-in rule on 8, while Bonferroni exceeds on none. At 1.0 percent each of the three rules exceeds on 1 split, and at 1.5 and 2.0 percent none does. (c) Pairs of horizontal bars, the answer-level and the problem-count upper bound on risk in percent: PRM 1.41 and 2.14, 8B probe 0.91 and 1.62, 8B re-solve 1.09 and 1.89, adaptive vote 0.91 and 1.64, unanimous 3-vote 0.47 and 1.20, small-first cascade 0.73 and 1.45.
(a) Predicted minus realised coverage of two composed cascades on three answer pools, from a probe of 100 labels. (b) Splits, of 15, whose certified selection exceeds the target on test under three testing rules. (c) Nominal answer-level and problem-count 95% upper bounds on selective risk.

At the strictest target, an uncorrected \(p\)-value and a plug-in rule that admits any schedule whose calibration risk is at most the target both certify every split but exceed the target on test, whereas the flat Bonferroni correction certifies a third of the splits and exceeds on none. Answers to the same problem are dependent, with an intraclass correlation of 0.45 for the pool's error indicator, and counting distinct kept problems widens every reported bound while preserving their ordering.

Prices Across Populations and Cost Axes

Both rates shift across actions on two further mathematics benchmarks and on program synthesis. On every population, each action agrees with correct answers clearly more often than with wrong ones, so a price refitted on a new population is still a usable input to the same composition.

Three panels. (a) Completeness per draw in percent for the 8B probe, the 8B re-solve and the 1.7B re-solve: 70.1, 80.2 and 79.5 on MATH, 72.9, 84.2 and 72.0 on Minerva, 97.6, 94.0 and 93.9 on GSM8K, and 97.4 and 75.4 for the probe and the 8B regeneration on programs. (b) Leak per draw in percent: 10.8, 9.0 and 10.6 on MATH, 30.5, 43.5 and 28.0 on Minerva, 16.3, 23.3 and 22.3 on GSM8K, and 2.2 and 3.1 on programs. (c) Cost per action relative to an 8B re-solve, in weighted tokens and in estimated seconds: about 0.04 and 0.07 for the 1.7B probe, 0.10 and 0.10 for the 4B probe, 0.19 and 0.14 for the 8B probe, and 0.19 and 0.49 for the 1.7B re-solve.
(a) Completeness and (b) leak per draw of three actions on four answer populations, where on programs the re-solve regenerates the program. (c) Parameter-weighted tokens and accelerator-time estimates from measured throughput, relative to an 8B re-solve.

The parameter-weighted proxy overstates the small model's advantage and reverses the ordering of a 1.7B re-solve and an 8B probe. The pooled small-first cascade is estimated to use about a third of the adaptive 8B vote's weighted tokens and about 30% less accelerator time.

Prices depend on the answer population and cost model, so applying PriceCheck to further model families, tasks or deployment settings calls for refitting the prices and recalibrating on representative data.

BibTeX

@misc{lee2026pricecheck,
  title = {Risk-Controlled Selective LLM Answering by Pricing Label-Free Checks},
  author = {Dongyub Jude Lee and Jungseob Lee and Chanjun Park and Hyeonseok Moon and Heuiseok Lim},
  year = {2026},
  journal = {arXiv preprint},
  eprint = {2609.37493},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.37493},
}