Risk-Controlled Selective LLM Answering by Pricing Label-Free Checks
arXiv preprint
1Zoom Communications 2Korea University 3Soongsil University 4Sookmyung Women's University
*Equal contribution †Corresponding authors
Abstract
Serving an answer from a large language model requires deciding when to abstain, yet a verifier's ranking accuracy alone does not determine the error rate among served answers. We introduce PriceCheck, which builds a compact family of decision rules from label-free checks such as re-solving a problem. Each check has a price: its agreement rates on correct and incorrect answers and its cost per run. Prices fitted on a small, class-enriched labelled set compose into predictions of a schedule's coverage and cost, guiding which checks to run and when to stop. A calibration test then selects a schedule at a stated selective-risk target. In mathematics, the selected schedules serve 76.1% of answers on average and keep held-out selective risk below 1.5% on all 15 splits. Under the shared testing protocol, PriceCheck serves more answers at that target than reward models, a prompted judge, the generator's confidence and a trained correctness classifier. At matched coverage, it keeps the fewest wrong answers among these scorers. Across 118 diagnostic schedules, price-based coverage predictions have a rank correlation of 0.97 with observed coverage. These results show that choosing how checks are combined and stopped matters alongside how well a verifier ranks answers. Code is available at https://
PriceCheck
A check reads a candidate answer, but not its label, and returns a binary verdict on whether it agrees with that answer. A schedule \(S\) uses one or more checks to serve an answer or abstain. Its coverage is the share of answers served, and its selective risk \(R(S)\) is the error rate among them. Given a target \(\alpha\) and a confidence budget \(\delta\), the objective is to select and certify a schedule for which \(R(S) \le \alpha\) holds with probability at least \(1 - \delta\) over the draw of calibration data, at the largest coverage obtainable.
-
Step 1
Price each check
The price of a check is two measured rates and a unit cost. Completeness \(p_1\) is the per-draw rate at which the check agrees with a correct answer, and leak \(p_0\) the rate at which it agrees with a wrong one.
\[\begin{aligned} p_1 &= \mathbb{P}(V = 1 \mid Y = 1) \\ p_0 &= \mathbb{P}(V = 1 \mid Y = 0) \end{aligned}\]
Here \(V\) is the verdict of one draw and \(Y\) the correctness of the answer. Both rates are fitted on a small labelled pricing set.
-
Step 2
Compose prices into a schedule
Prices compose into predictions of a schedule's coverage, selective risk and cost. With \(a_y\) the fitted acceptance probability for class \(y\) and \(\pi\) the correct-answer fraction of the pricing set, predicted coverage and risk are
\[\begin{aligned} \hat{c} &= \pi\, a_1 + (1-\pi)\, a_0 \\ \hat{r} &= \frac{(1-\pi)\, a_0}{\hat{c}} \end{aligned}\]
Cost composes in the same way, as each stage is paid only by the answers that reach it.
-
Step 3
Certify one schedule
Each schedule in a family \(G\), fixed independently of the i.i.d. calibration answers, gets an exact one-sided binomial \(p\)-value against \(H_0\colon R(S) > \alpha\). It is certified when it serves \(n > 0\) calibration answers, \(e\) of them wrong, and
\[F(e;\, n,\, \alpha) \le \delta / |G|\]
with \(F\) the binomial cumulative distribution function, \(\delta = 0.05\) and \(|G| = 23\). The default selector takes the certified schedule with the largest calibration coverage.
Schedules over Priced Checks
The paper calls a check an action and prices seven of them. A backward probe has a reverse model write the reasoning that leads to the proposed answer, and a forward model then recovers an answer from that trace. A re-solve vote draws a fresh solution from the frozen forward model. Two generator-confidence actions read the generator's probability of Yes when asked whether the proposed answer is correct, and the mean log-probability of the stored solution.
| Backward probe | Re-solve vote | Generator confidence | |||||
|---|---|---|---|---|---|---|---|
| 8B | 4B | 1.7B | 8B | 1.7B | p(True) | Log-prob. | |
| Rate per draw (%) | |||||||
| Completeness \(p_1\) | 70.1 | 73.1 | 63.9 | 80.2 | 79.5 | 72.5 | 61.9 |
| Leak \(p_0\) | 10.8 | 13.8 | 12.5 | 9.0 | 10.6 | 54.5 | 30.3 |
| Cost per draw (tokens) | |||||||
| Generated | 1034 | 1034 | 1034 | 5384 | 4801 | 0 | 0 |
| Weighted | 1034 | 517 | 217 | 5384 | 1008 | 0 | 0 |
Scroll sideways to see every column.
Price of each action on the frozen pool: completeness \(p_1\) and leak \(p_0\) per draw, and unit cost per draw in generated and in parameter-weighted tokens. The two confidence actions re-score the stored solution, so they generate nothing.
Schedules are built from four primitives. A threshold serves when the agreement count of one backward probe reaches \(t\), and unanimity when the first \(N\) re-solve votes all agree. An adaptive race draws re-solve votes until it reaches \(k\) agreements, where it serves, or \(j\) disagreements, where it abstains. A cascade runs a cheap action first and sends its undecided band to the race. The deployed family is a fixed grid of 23 schedules that no certification run alters.
Main Results
Five generators, all fine-tuned from one 8B backbone, each attempted the 500 problems of MATH-500, and the evaluation uses 2326 of their answers with a base error of 4.26%. PriceCheck is compared with a prompted judge, a trained process reward model (PRM), Math-Shepherd, and an outcome reward model (ORM) under the shared testing protocol across fifteen splits. For a fair comparison, each rival scorer receives a family of selective rules of the same shape and comparable size, and the comparisons apply the test at its nominal level and assess realised held-out risk.
| Method | \(\alpha\) = 0.5% | 0.75% | 1.0% | 1.25% | 1.5% | 2.0% | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Spl. | Cov. | Spl. | Cov. | Spl. | Cov. | Spl. | Cov. | Spl. | Cov. | Spl. | Cov. | |
| PriceCheck (ours) | 5 | 23.20 | 12 | 57.00 | 15 | 72.11 | 15 | 73.85 | 15 | 76.09 | 15 | 78.86 |
| Prompted verifier | ||||||||||||
| Judge | 0 | 0 | 5 | 16.74 | 15 | 52.43 | 15 | 62.32 | 15 | 69.50 | 15 | 78.17 |
| Judge† | 0 | 0 | 5 | 20.09 | 15 | 61.08 | 15 | 67.99 | 15 | 71.91 | 15 | 80.81 |
| Trained verifiers | ||||||||||||
| PRM | 0 | 0 | 1 | 3.31 | 5 | 16.71 | 8 | 26.77 | 15 | 55.90 | 15 | 68.14 |
| PRM† | 0 | 0 | 1 | 3.90 | 8 | 32.09 | 15 | 62.46 | 15 | 66.55 | 15 | 74.20 |
| Math-Shepherd | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 3.56 | 5 | 17.52 |
| Math-Shepherd† | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 2 | 7.61 | 6 | 23.80 |
| ORM | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 3.47 | 3 | 10.23 | 11 | 38.63 |
| ORM† | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 2 | 7.66 | 9 | 34.94 |
Scroll sideways to see every column.
Certification under a shared protocol. At each target \(\alpha\), splits certifying of 15 and coverage over all 15 counting non-certifying splits as zero. A dagger marks 24 levels, not six. Bold is best. The last column of the paper's table, wrong answers kept at matched coverage, is given in the text below.
PriceCheck certifies on all fifteen splits at the primary 1.5% target, with the highest mean held-out coverage, 76.09%, and no test-risk exceedance. Its coverage advantage over the best rival grows from 4.18 points at the primary target to 36.91 points at 0.75%, and at the strictest 0.5% target only PriceCheck certifies any split. The judge's finer family leads at the looser 2% target.
At matched coverage, each rival's test answers are ranked and exactly as many are kept as the certified schedule kept on that split. Every rival scorer then keeps more wrong answers than PriceCheck's 44. The count is 60 for the judge, the strongest ranker, 97 for the PRM, 164 for Math-Shepherd and 195 for the ORM. Neither generator-confidence score certifies on any split at any target, and a 1.7B correctness classifier fitted to the calibration labels certifies on only three of the fifteen splits even at its widest context window.
Across Split Schemes
The fifteen splits are five source-held-out folds, each reserving one generator, and ten question-disjoint halves that reserve problems. Certified coverage is higher on the folds than on the halves. The default selector certifies an 8B re-solve vote on every split, while a cost-minimising selector instead picks a small-first cascade on three of the five folds, where its selections average 77.7% coverage at under a third of the adaptive vote's weighted cost.
Panel (c) certifies the plain fraction of stored re-solves agreeing with an answer as a scalar score under the same protocol. With three 8B re-solves, which generate about as many tokens as the certified schedule, this score serves less at the primary target, and it reaches the schedule's coverage only at sixteen votes.
Prices Predict Schedule Coverage
Prices fitted on 100 class-enriched answers track cascade coverage across three pools. Over 118 diagnostic schedules, price-based coverage predictions have a rank correlation of 0.97 with observed coverage, and the mean absolute coverage error is 2.97 points over 70 cascades. Risk prediction has a mean error of about one point but is less reliable across schedules, so PriceCheck tests schedules on calibration data rather than using predicted risk as a certificate.
At the strictest target, an uncorrected \(p\)-value and a plug-in rule that admits any schedule whose calibration risk is at most the target both certify every split but exceed the target on test, whereas the flat Bonferroni correction certifies a third of the splits and exceeds on none. Answers to the same problem are dependent, with an intraclass correlation of 0.45 for the pool's error indicator, and counting distinct kept problems widens every reported bound while preserving their ordering.
Prices Across Populations and Cost Axes
Both rates shift across actions on two further mathematics benchmarks and on program synthesis. On every population, each action agrees with correct answers clearly more often than with wrong ones, so a price refitted on a new population is still a usable input to the same composition.
The parameter-weighted proxy overstates the small model's advantage and reverses the ordering of a 1.7B re-solve and an 8B probe. The pooled small-first cascade is estimated to use about a third of the adaptive 8B vote's weighted tokens and about 30% less accelerator time.
Prices depend on the answer population and cost model, so applying PriceCheck to further model families, tasks or deployment settings calls for refitting the prices and recalibrating on representative data.
BibTeX
@misc{lee2026pricecheck,
title = {Risk-Controlled Selective LLM Answering by Pricing Label-Free Checks},
author = {Dongyub Jude Lee and Jungseob Lee and Chanjun Park and Hyeonseok Moon and Heuiseok Lim},
year = {2026},
journal = {arXiv preprint},
eprint = {2609.37493},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.37493},
}