Jungseob Lee Publications

DART: Draft-Agreement Routing for Training-Free Adaptive Thinking Budgets in Hybrid Reasoning Models

Findings of EMNLP 2026

Jungseob Lee1, Seongtae Hong1, Seungjun Lee2, Jaehyung Seo3, Junyoung Son1, Sugyeong Eo4, Chanjun Park5, Hyeongju Park6, Hyeonseok Moon1*, Heuiseok Lim1*

1Korea University 242dot 3Konkuk University 4Yonsei University 5Soongsil University 6Kumoh National Institute of Technology

*Corresponding authors

A training-free router for hybrid reasoning models: cheap no-think drafts decide when to think, and how much.

DART pipeline. Stage 1, SC-Route, draws two no-think drafts for a query and checks whether their extracted answers agree. When they agree, the draft answer is accepted with no thinking. When they disagree, Stage 2 maps draft entropy to a thinking budget, caps the thinking trace at that budget, and produces the answer in a separate completion.
Overview of DART. Stage 1 draws \(K = 2\) no-think drafts and accepts the unanimous answer when they agree. On disagreement, Stage 2 maps draft entropy to a query-specific thinking budget and produces the answer in a separate completion.

PromptEvaluate \((1+2i)6-3i\).

  • DART (\(K{=}2\)) 10.5×over AT 332tokens 0 thinking tokens10.8× fewer tokens
    12.29 s
  • AT always think 1×baseline 347tokens 3,407 thinking tokens1 thinking call
    12.29 s

Qwen3-8B on one RTX A6000

Abstract

Hybrid reasoning models can answer directly or spend extra tokens on extended thinking. A practical router should choose between these modes for each query, so easy problems avoid unnecessary reasoning and hard problems receive enough budget to finish the answer. Existing routers move in this direction, but they typically require labeled training data or fix thinking budgets up front, ignoring answer-level evidence from the model itself. We introduce DART, a training-free routing framework that samples two cheap no-think drafts, accepts direct answering when the drafts agree, and predicts a thinking budget from draft entropy when they disagree. Across the main comparisons, DART preserves or improves always-thinking accuracy in most settings while reducing thinking-token use. Accuracy improves by up to +9.0 points on Olympiad-level math and by up to +22.5 points on code under execution-based equivalence, while thinking-token use drops by 32–73%. The Stage 1 signal extends across model scales (0.6B–32B), model families, and API-only hosted settings, with no labeled data and no gradient updates required. Our code is available at https://github.com/js-lee-AI/DART.

Method

DART uses cheap NoThink drafts as a query-level difficulty probe. The pipeline requires no labeled difficulty data or gradient updates, and it separates answer generation from the thinking trace to avoid single-call truncation artifacts.

  1. Stage 1

    SC-Route: accept when the drafts agree

    Sample \(K\) independent completions in NoThink mode and extract an answer from each. If all extracted answers are equivalent, accept the draft answer. Otherwise, route the query to Think mode.

    The equivalence function is the only domain-specific component: string normalization for math, and sandboxed execution against each problem's reference tests for code. Stage 1 has one hyperparameter (\(K\), the number of drafts) and no thresholds. The main experiments use strict unanimity at \(K = 2\).

  2. Stage 2

    Budget prediction: entropy sets the cap

    Stage 2 is an evaluation-time protocol applied only to queries Stage 1 routes to Think mode. Draft entropy \(H(q)\), the mean token-level entropy over the drafts, sets a query-specific thinking budget.

    \[\hat{B}(q) = \gamma \cdot f\bigl(H(q)\bigr)\]

    Here \(f\) is an isotonic regression fit without accuracy labels on a separate 100-problem probe run, 71 of whose problems also appear in the MATH-500 evaluation set, and \(\gamma = 1.5\) is a safety margin. For models exposing a stop-thinking control, such as Qwen3, a </think> token is injected when the trace reaches the budget, and the answer is generated in a separate completion call.

Main Results

DART is compared with always-thinking (AT), which emits the thinking trace and the final answer in separate completion calls, and with no-think (NT), which keeps thinking off for all queries. Across the observed point estimates, DART matches or exceeds AT on 13 of 14 model–benchmark pairs and reduces thinking-token use by 32–73%. The single miss is Qwen3-8B on OlympiadBench, where DART trails AT by 1.7 points while still reducing thinking tokens by 45%.

BenchmarkAccuracy (%)Efficiency
NTATSC-RouteDART (vs AT)Think ↓Route%
Qwen3-8B
MATH-50076.685.687.688.2(+2.6)67%78.0
OlympiadBench49.871.570.469.8(−1.7)45%52.1
HumanEval60.459.178.778.7(+19.6)55%54.9
MBPP59.564.267.568.9(+4.7)46%57.6
Qwen3-14B
MATH-50081.287.687.687.6(+0.0)73%83.4
OlympiadBench51.553.062.062.0(+9.0)51%58.0
HumanEval71.366.578.778.7(+12.2)60%68.9
MBPP64.664.668.168.1(+3.5)51%63.0
Qwen3-32B
MATH-50082.286.288.588.5(+2.3)69%80.0
OlympiadBench50.554.058.558.5(+4.5)52%59.5
HumanEval79.972.695.195.1(+22.5)63%76.8
MBPP65.865.871.271.2(+5.4)51%63.0
DeepSeek-V3.2*
MATH-50084.888.490.692.6(+4.2)56%85.0
OlympiadBench63.666.167.169.1(+3.0)32%60.1

Scroll sideways to see every column.

Accuracy (in %) and thinking-token efficiency on Qwen3-8B/14B/32B and DeepSeek-V3.2. Accuracy columns report NT, AT, SC-Route (Stage 1 routing only, with disagreement queries falling back to AT), and the full DART pipeline (with vs-AT delta in green/red). Efficiency columns report Think ↓, the thinking-token reduction vs AT, and Route%, the rate at which Stage 1 accepts unanimous drafts. Shading marks DART matching or exceeding AT, and bold marks the better of DART and AT in each row. *DeepSeek-V3.2 results use the hosted API.

Rows are single-run and cover the full benchmark sizes, except that Qwen3-32B MATH-500 covers 348 of 500 problems and OlympiadBench covers 200 of 572 problems at Qwen3-14B and Qwen3-32B.

On code, DART improves over AT by up to +22.5 points while reducing thinking-token use by 46–63%. Code agreement is execution-based: both drafts are run against each problem's reference tests, the same tests that score the final answer, so an accepted code query passes by construction. The large HumanEval gains arise because AT underperforms NT on short-form code, and Stage 1 keeps high-agreement code queries on the direct-answer path.

Token and Latency Efficiency

On MATH-500 with Qwen3-8B, DART averages 3.5K total generated tokens compared with 5.4K for AT and gains 2.6 accuracy points, a 35% total-token reduction that counts drafts, thinking, and answers. On HumanEval with the same model, DART reduces thinking tokens by 55% while gaining 19.6 points.

Token savings also reduce wall-clock latency. With Qwen3-8B served locally on a single A100-80GB node, DART averages 29.5 seconds per MATH-500 query, compared with 67.6 seconds for AT. The 2.3-fold speedup comes from the 78% of queries accepted after parallel NoThink drafts, which finish in 14.6 seconds on average and avoid a Think pass. These timings describe one serving configuration.

Left: MATH-500 accuracy against mean thinking tokens per query. The one-phase always-thinking sweep rises from about 14% at just under 2K tokens to about 72% at just over 5K tokens, two-phase always-thinking sits near 86% at about 5.3K tokens, and DART sits near 88% at about 1.7K tokens. Right: stacked bars of truncated, incorrect, and correct responses for one-phase always-thinking at 2K to 32K budgets and for DART. Truncation falls from 86% at 2K to about 1% at 32K, where 72% are correct. DART has 88% correct, 12% incorrect, and no truncation.
Efficiency Pareto (left) and response-level failure-mode breakdown (right) for Qwen3-8B on MATH-500. Both panels place each method at the thinking tokens it actually emits. The one-phase sweep uses the paper's diagnostic grader, while the two-phase AT and DART points use the main protocol of the table above.

Difficulty Alignment

Stage 1 accepts fewer queries as annotated difficulty increases while preserving high precision among accepted answers. On MATH-500 with Qwen3-8B, the Stage 1 accept rate decreases monotonically from 97.7% at Level 1 to 59.7% at Level 5, while accepted-answer precision remains between 83.8% and 95.5% across all five levels.

Aggregated over all Qwen3-8B queries, accepted answers are 90.8% correct on MATH-500 and 81.9% on OlympiadBench, 14 and 32 points above the unconditional NoThink accuracies of 76.6 and 49.8. Across the settings of the main table and a Qwen3 size sweep, draft agreement remains positively associated with AT correctness, with a positive and significant point-biserial correlation across two model lineages, two architectures, and scales from 0.6B to 32B. Accuracy parity depends on model capability, and DART trails AT by 2.0–7.5 points on MATH-500 at 0.6B–4B.

Left: stacked bars of routing share by MATH-500 difficulty level. The accepted share is 98, 89, 85, 77, and 60 percent at levels 1 to 5, and the rest is escalated to thinking. Right: accept precision by level, 93, 92, 96, 90, and 84 percent at levels 1 to 5.
Difficulty-stratified Stage-1 behavior on MATH-500 (Qwen3-8B). The left panel shows the no-think accept share and the share escalated to thinking across difficulty levels 1–5. The right panel shows precision among accepted no-think answers at each level.

BibTeX

@article{lee2026dart,
  title={DART: Draft-Agreement Routing for Training-Free Adaptive Thinking Budgets in Hybrid Reasoning Models},
  author={Lee, Jungseob and Hong, Seongtae and Lee, Seungjun and Seo, Jaehyung and Son, Junyoung and Eo, Sugyeong and Park, Chanjun and Park, Hyeongju and Moon, Hyeonseok and Lim, Heuiseok},
  journal={arXiv preprint arXiv:2606.23181},
  year={2026}
}