DART: Draft-Agreement Routing for Training-Free Adaptive Thinking Budgets in Hybrid Reasoning Models
Findings of EMNLP 2026
1Korea University 242dot 3Konkuk University 4Yonsei University 5Soongsil University 6Kumoh National Institute of Technology
*Corresponding authors
PromptEvaluate \((1+2i)6-3i\).
-
DART (\(K{=}2\)) 10.5×over AT 332tokens12.29 s
-
AT always think 1×baseline 347tokens12.29 s
Qwen3-8B on one RTX A6000
Abstract
Hybrid reasoning models can answer directly or spend extra tokens on extended thinking. A practical router should choose between these modes for each query, so easy problems avoid unnecessary reasoning and hard problems receive enough budget to finish the answer. Existing routers move in this direction, but they typically require labeled training data or fix thinking budgets up front, ignoring answer-level evidence from the model itself. We introduce DART, a training-free routing framework that samples two cheap no-think drafts, accepts direct answering when the drafts agree, and predicts a thinking budget from draft entropy when they disagree. Across the main comparisons, DART preserves or improves always-thinking accuracy in most settings while reducing thinking-token use. Accuracy improves by up to +9.0 points on Olympiad-level math and by up to +22.5 points on code under execution-based equivalence, while thinking-token use drops by 32–73%. The Stage 1 signal extends across model scales (0.6B–32B), model families, and API-only hosted settings, with no labeled data and no gradient updates required. Our code is available at https://
Method
DART uses cheap NoThink drafts as a query-level difficulty probe. The pipeline requires no labeled difficulty data or gradient updates, and it separates answer generation from the thinking trace to avoid single-call truncation artifacts.
-
Stage 1
SC-Route: accept when the drafts agree
Sample \(K\) independent completions in NoThink mode and extract an answer from each. If all extracted answers are equivalent, accept the draft answer. Otherwise, route the query to Think mode.
The equivalence function is the only domain-specific component: string normalization for math, and sandboxed execution against each problem's reference tests for code. Stage 1 has one hyperparameter (\(K\), the number of drafts) and no thresholds. The main experiments use strict unanimity at \(K = 2\).
-
Stage 2
Budget prediction: entropy sets the cap
Stage 2 is an evaluation-time protocol applied only to queries Stage 1 routes to Think mode. Draft entropy \(H(q)\), the mean token-level entropy over the drafts, sets a query-specific thinking budget.
\[\hat{B}(q) = \gamma \cdot f\bigl(H(q)\bigr)\]
Here \(f\) is an isotonic regression fit without accuracy labels on a separate 100-problem probe run, 71 of whose problems also appear in the MATH-500 evaluation set, and \(\gamma = 1.5\) is a safety margin. For models exposing a stop-thinking control, such as Qwen3, a
</think>token is injected when the trace reaches the budget, and the answer is generated in a separate completion call.
Main Results
DART is compared with always-thinking (AT), which emits the thinking trace and the final answer in separate completion calls, and with no-think (NT), which keeps thinking off for all queries. Across the observed point estimates, DART matches or exceeds AT on 13 of 14 model–benchmark pairs and reduces thinking-token use by 32–73%. The single miss is Qwen3-8B on OlympiadBench, where DART trails AT by 1.7 points while still reducing thinking tokens by 45%.
| Benchmark | Accuracy (%) | Efficiency | ||||
|---|---|---|---|---|---|---|
| NT | AT | SC-Route | DART (vs AT) | Think ↓ | Route% | |
| Qwen3-8B | ||||||
| MATH-500 | 76.6 | 85.6 | 87.6 | 88.2(+2.6) | 67% | 78.0 |
| OlympiadBench | 49.8 | 71.5 | 70.4 | 69.8(−1.7) | 45% | 52.1 |
| HumanEval | 60.4 | 59.1 | 78.7 | 78.7(+19.6) | 55% | 54.9 |
| MBPP | 59.5 | 64.2 | 67.5 | 68.9(+4.7) | 46% | 57.6 |
| Qwen3-14B | ||||||
| MATH-500 | 81.2 | 87.6 | 87.6 | 87.6(+0.0) | 73% | 83.4 |
| OlympiadBench | 51.5 | 53.0 | 62.0 | 62.0(+9.0) | 51% | 58.0 |
| HumanEval | 71.3 | 66.5 | 78.7 | 78.7(+12.2) | 60% | 68.9 |
| MBPP | 64.6 | 64.6 | 68.1 | 68.1(+3.5) | 51% | 63.0 |
| Qwen3-32B | ||||||
| MATH-500 | 82.2 | 86.2 | 88.5 | 88.5(+2.3) | 69% | 80.0 |
| OlympiadBench | 50.5 | 54.0 | 58.5 | 58.5(+4.5) | 52% | 59.5 |
| HumanEval | 79.9 | 72.6 | 95.1 | 95.1(+22.5) | 63% | 76.8 |
| MBPP | 65.8 | 65.8 | 71.2 | 71.2(+5.4) | 51% | 63.0 |
| DeepSeek-V3.2* | ||||||
| MATH-500 | 84.8 | 88.4 | 90.6 | 92.6(+4.2) | 56% | 85.0 |
| OlympiadBench | 63.6 | 66.1 | 67.1 | 69.1(+3.0) | 32% | 60.1 |
Scroll sideways to see every column.
Accuracy (in %) and thinking-token efficiency on Qwen3-8B/14B/32B and DeepSeek-V3.2. Accuracy columns report NT, AT, SC-Route (Stage 1 routing only, with disagreement queries falling back to AT), and the full DART pipeline (with vs-AT delta in green/red). Efficiency columns report Think ↓, the thinking-token reduction vs AT, and Route%, the rate at which Stage 1 accepts unanimous drafts. Shading marks DART matching or exceeding AT, and bold marks the better of DART and AT in each row. *DeepSeek-V3.2 results use the hosted API.
Rows are single-run and cover the full benchmark sizes, except that Qwen3-32B MATH-500 covers 348 of 500 problems and OlympiadBench covers 200 of 572 problems at Qwen3-14B and Qwen3-32B.
On code, DART improves over AT by up to +22.5 points while reducing thinking-token use by 46–63%. Code agreement is execution-based: both drafts are run against each problem's reference tests, the same tests that score the final answer, so an accepted code query passes by construction. The large HumanEval gains arise because AT underperforms NT on short-form code, and Stage 1 keeps high-agreement code queries on the direct-answer path.
Token and Latency Efficiency
On MATH-500 with Qwen3-8B, DART averages 3.5K total generated tokens compared with 5.4K for AT and gains 2.6 accuracy points, a 35% total-token reduction that counts drafts, thinking, and answers. On HumanEval with the same model, DART reduces thinking tokens by 55% while gaining 19.6 points.
Token savings also reduce wall-clock latency. With Qwen3-8B served locally on a single A100-80GB node, DART averages 29.5 seconds per MATH-500 query, compared with 67.6 seconds for AT. The 2.3-fold speedup comes from the 78% of queries accepted after parallel NoThink drafts, which finish in 14.6 seconds on average and avoid a Think pass. These timings describe one serving configuration.
Difficulty Alignment
Stage 1 accepts fewer queries as annotated difficulty increases while preserving high precision among accepted answers. On MATH-500 with Qwen3-8B, the Stage 1 accept rate decreases monotonically from 97.7% at Level 1 to 59.7% at Level 5, while accepted-answer precision remains between 83.8% and 95.5% across all five levels.
Aggregated over all Qwen3-8B queries, accepted answers are 90.8% correct on MATH-500 and 81.9% on OlympiadBench, 14 and 32 points above the unconditional NoThink accuracies of 76.6 and 49.8. Across the settings of the main table and a Qwen3 size sweep, draft agreement remains positively associated with AT correctness, with a positive and significant point-biserial correlation across two model lineages, two architectures, and scales from 0.6B to 32B. Accuracy parity depends on model capability, and DART trails AT by 2.0–7.5 points on MATH-500 at 0.6B–4B.
BibTeX
@article{lee2026dart,
title={DART: Draft-Agreement Routing for Training-Free Adaptive Thinking Budgets in Hybrid Reasoning Models},
author={Lee, Jungseob and Hong, Seongtae and Lee, Seungjun and Seo, Jaehyung and Son, Junyoung and Eo, Sugyeong and Park, Chanjun and Park, Hyeongju and Moon, Hyeonseok and Lim, Heuiseok},
journal={arXiv preprint arXiv:2606.23181},
year={2026}
}