ThinkFuse: Trajectory-Aware Test-Time Fusion for Small Reasoning Models
Findings of EMNLP 2026
1Department of Computer Science and Engineering, Korea University 2Human-inspired AI Research 3Department of Computer Science and Engineering, Konkuk University
*Equal contribution †Corresponding author
QuestionA regular pentagon is rotated counterclockwise about its center. What is the minimum number of degrees it must be rotated until it coincides with its original position?
Answer72
-
ThinkFuse
72correct
1,941 tokens329 frames -
Cool-Fusion –no answer1,936 tokens329 frames
-
Qwen3-4B alone 360wrong1,936 tokens329 frames
Qwen3-4B with Ministral-3B-R, blue is the text ThinkFuse took from the auxiliary model
Abstract
Small reasoning models (SRMs) have shown strong performance on complex reasoning tasks by generating extended chain-of-thought trajectories, but they often fail to recover once their reasoning enters an erroneous path. Existing test-time fusion methods rely on local fusion signals to determine when to trigger fusion, which can be misled by transient uncertainty fluctuations and may reinforce unstable reasoning trajectories. We propose ThinkFuse, a training-free test-time fusion framework that selectively intervenes in unreliable reasoning segments. ThinkFuse compares segment-level uncertainty shifts with trajectory-level uncertainty trends to identify unstable reasoning points and fuse auxiliary reasoning paths into the primary model's trajectory. Extensive experiments demonstrate that ThinkFuse outperforms baselines on mathematical and knowledge-intensive reasoning benchmarks, with consistent gains across model-family combinations, and remains robust with a smaller primary model. Our analysis shows that ThinkFuse requires fewer fusion triggers and generates fewer tokens, highlighting the efficiency of selective triggering. Our code is available at https://
Method
ThinkFuse prompts the primary reasoning model \(M_p\) to generate a reasoning path and monitors its uncertainty at the segment level, capturing both the variation within a segment and the overall progress of the cumulative trajectory. Once the segment-level uncertainty exceeds an adaptive threshold, ThinkFuse evaluates the primary and auxiliary segments by primary-model perplexity under the current context and extends the reasoning trajectory with the more compatible segment.
-
Step 1
Token uncertainty over fixed-length segments
ThinkFuse monitors uncertainty at the segment level rather than relying directly on isolated token-level signals, which can be sensitive to local lexical or discourse fluctuations. Segments have a fixed length of \(T\) tokens. For each token of the segment \(s_k^p\) that the primary model \(M_p\) generates from the context \(c_k\), it computes normalized Shannon entropy over the top-\(r\) log-probability distribution:
\[u_{k,i}^{p} = -\frac{\sum_{v \in \mathcal{V}_{k,i}} \tilde{p}_{k,i}(v) \log \tilde{p}_{k,i}(v)}{\log \lvert \mathcal{V}_{k,i} \rvert}\]
Here \(\mathcal{V}_{k,i}\) is the set of the \(r\) highest-probability candidate tokens at position \(i\), and \(\tilde{p}_{k,i}\) is the distribution renormalized over that set.
-
Step 2
Segment uncertainty with EWCA
Exponentially Weighted Causal Aggregation (EWCA) turns the token uncertainties into one segment-level signal. Each token uncertainty is weighted by its relative position and by the uncertainty accumulated over the preceding tokens:
\[w_{k,i} = \exp(\alpha \cdot i) \cdot \prod_{j<i} \bigl(1 + \beta \cdot u_{k,j}^{p}\bigr)\]
where \(\alpha\) controls later-token emphasis and \(\beta\) controls uncertainty accumulation. The segment-level uncertainty of the \(k\)-th segment is
\[U_k^p = \frac{\sum_{i=1}^{T_k} w_{k,i}\, u_{k,i}^p}{\sum_{i=1}^{T_k} w_{k,i}}\]
-
Step 3
Adaptive threshold with a soft budget
The same uncertainty value can have different implications depending on the recent reasoning trajectory. The mean uncertainty \(\mu_k\) and deviation scale \(\sigma_k\) of the trajectory are updated by an exponential moving average with decay factor \(\eta\), and they set the threshold
\[\begin{aligned} \mu_k &= (1-\eta)\,\mu_{k-1} + \eta\, U_k^p \\ \sigma_k &= (1-\eta)\,\sigma_{k-1} + \eta\, \lvert U_k^p - \mu_{k-1} \rvert \\ \theta_k &= (\mu_{k-1} + \tau \cdot \sigma_{k-1}) \cdot (1 + \lambda \cdot R_{k-1}^2) \end{aligned}\]
where \(\tau\) controls the deviation margin and \(R_{k-1}\) is the cumulative fusion ratio, scaled by the soft budget parameter \(\lambda\). The quadratic term raises the threshold as fusions accumulate, and fusion is triggered only when \(U_k^p \ge \theta_k\).
-
Step 4
Trajectory compatibility scoring
When the trigger activates, the auxiliary reasoning model \(M_a\) generates an alternative segment \(s_k^a\) from the same context. The primary model scores both candidates as a trajectory-compatibility scorer, \(\ell_k(s) = \mathrm{PPL}_{M_p}(s \mid c_k)\), and the selected segment is
\[s_k^\star = \begin{cases} s_k^p, & U_k^p < \theta_k \\ \arg\min\limits_{s \in \{s_k^p,\, s_k^a\}} \ell_k(s), & U_k^p \ge \theta_k \end{cases}\]
It is appended to the reasoning context, \(c_{k+1} = c_k \oplus s_k^\star\), and the primary model continues generation from the updated context until the reasoning trajectory reaches a terminal state.
Monitoring and fusion apply only inside the explicit thinking stage, and the auxiliary candidate is generated in parallel with the primary model whenever applicable. The default setting uses segments of \(T = 8\) tokens, top-5 token entropy, EWCA with \((\alpha, \beta) = (0.1, 0.5)\), an EMA update rate \(\eta = 0.05\), a tolerance \(\tau = 1.0\), and a soft-budget coefficient \(\lambda = 50\).
Main Results
The comparison covers standalone models, self-consistency with Qwen3-4B, the test-time fusion methods Cool-Fusion and AdaFuse, and ThinkFuse with cross-family and same-family model pairs. All methods share a 16,384-token generation budget. Each method is evaluated on the same fixed random subsample of 200 examples per benchmark, and AIME 2024 (30 problems) and GPQA-Diamond (198 problems) are evaluated in full. Ministral stands for Ministral-3B-R and EXAONE-4.0 for EXAONE-4.0-1.2B.
With Qwen3-4B as the primary model, ThinkFuse improves over its standalone performance on every benchmark, whether the auxiliary is a cross-family model or a smaller model from the same family. In particular, Qwen3-4B × Ministral achieves the best result on four out of five benchmarks, with large gains on AIME24 (+20.0), GPQA (+10.6), and NQ-Open (+10.0).
| Method | MATH-500 | GSM8K | AIME24 | GPQA | NQ-Open |
|---|---|---|---|---|---|
| Standalone models | |||||
| Qwen3-4B | 89.0 | 92.0 | 53.3 | 54.0 | 29.5 |
| Qwen3-1.7B‡ | 83.0 | 90.0 | 33.3 | 39.4 | 28.0 |
| Ministral | 32.5 | 67.0 | 0.0† | 25.8 | 24.0 |
| EXAONE-4.0 | 57.0 | 85.0 | 3.3† | 38.9 | 22.0 |
| Test-time compute baseline: Qwen3-4B | |||||
| Self-consistency (\(K = 3\)) | 89.5(+0.5) | 94.5(+2.5) | 53.3(+0.0) | 51.0(−3.0) | 39.0(+9.5) |
| Self-consistency (\(K = 5\)) | 88.0(−1.0) | 94.5(+2.5) | 53.3(+0.0) | 50.0(−4.0) | 39.0(+9.5) |
| Test-time fusion baselines: Qwen3-4B × Ministral | |||||
| Cool-Fusion | 31.0(−58.0) | 66.0(−26.0) | 0.0†(−53.3) | 11.6(−42.4) | 0.5(−29.0) |
| AdaFuse | 45.5(−43.5) | 72.6(−19.4) | 33.3(−20.0) | 28.8(−25.2) | 38.5(+9.0) |
| ThinkFuse | |||||
| Qwen3-4B × Ministral | 90.0(+1.0) | 96.9(+4.9) | 73.3(+20.0) | 64.6(+10.6) | 39.5(+10.0) |
| Qwen3-4B × Qwen3-1.7B | 91.0(+2.0) | 94.0(+2.0) | 70.0(+16.7) | 55.1(+1.1) | 37.0(+7.5) |
| Qwen3-4B × Qwen3-4B | 75.0(−14.0) | 90.0(−2.0) | 40.0(−13.3) | 28.3(−25.7) | 38.0(+8.5) |
| Qwen3-1.7B × EXAONE-4.0 | 85.0(+2.0) | 90.5(+0.5) | 30.0(−3.3) | 53.5(+14.1) | 28.0(+0.0) |
Scroll sideways to see every column.
Main benchmark results. Accuracy (%) for standalone models, test-time baselines, and ThinkFuse across five benchmarks under a 16K-token budget. Values in parentheses indicate the signed accuracy difference, in percentage points, relative to the standalone performance of the primary model, with green indicating gains and red indicating degradation. The † symbol marks AIME24 results affected by length-limit or final-answer truncation. The ‡ symbol marks the standalone Qwen3-1.7B evaluated with the Qwen3 recommended top-\(k\) of 20, whereas all other rows use the shared decoding setting. Additional experimental details, including the decoding settings and remaining cross-family and same-family configurations, are presented in the appendix of the paper.
Existing test-time fusion baselines show the opposite trend. With the same Qwen3-4B × Ministral pair, Cool-Fusion drops sharply from 89.0 to 31.0 on MATH-500 and from 54.0 to 11.6 on GPQA, while AdaFuse also underperforms the standalone primary on four out of five benchmarks. Self-consistency provides gains on GSM8K and NQ-Open, but does not improve AIME24 and degrades GPQA. The redundant same-model variant, Qwen3-4B × Qwen3-4B, degrades performance on MATH-500, GSM8K, AIME24, and GPQA. These results suggest that fusion must be selective and trajectory-aware, since locally determined or redundant fusion can disrupt the primary model's reasoning process.
With the smaller primary model Qwen3-1.7B and EXAONE-4.0 as the auxiliary, ThinkFuse raises GPQA from 39.4 to 53.5, exceeding both standalone models. On the remaining benchmarks, where EXAONE-4.0 is substantially weaker than the primary, it gains +2.0 on MATH-500 and +0.5 on GSM8K, ties on NQ-Open, and differs by a single problem on AIME24. The standalone Qwen3-1.7B is evaluated with the Qwen3 recommended top-\(k\) of 20 while ThinkFuse retains the shared decoding setting, so this comparison favors the standalone baseline.
Fusion Dynamics and Downstream Performance
The fusion methods differ in how many auxiliary reasoning tokens they fuse and in how this changes the length of the thinking stage. Cool-Fusion fuses the most auxiliary tokens, yet its thinking stage is the shortest on four of the five benchmarks. This suggests that excessive fusion can overwrite or prematurely redirect the primary model's reasoning trajectory, explaining its large performance drops. AdaFuse uses fewer auxiliary tokens but still shortens the thinking stage on reasoning-intensive benchmarks such as AIME24 and GPQA, indicating that simply limiting auxiliary usage is insufficient for preserving reasoning quality.
ThinkFuse uses auxiliary tokens much more sparsely than Cool-Fusion. Combined with its accuracy gains, this analysis suggests that effective test-time fusion requires selective and trajectory-aware intervention, rather than frequent or locally determined auxiliary fusion.
Compute Efficiency
On the Qwen3-4B × Ministral pair, with accuracy averaged over the five benchmarks, ThinkFuse achieves the highest accuracy with a TFLOP budget comparable to AdaFuse and lower than that of Cool-Fusion, resulting in the highest accuracy per TFLOP. ThinkFuse exhibits a longer wall-clock time, yet the comparable TFLOPs indicate that this gap arises from the CPU-GPU data transfer during PPL scoring rather than from additional computation. The lower latency of the baselines comes at the cost of severe accuracy degradation, leaving ThinkFuse as the only fusion method that improves over the standalone primary. Compared with single-model decoding, serving an auxiliary model together with the primary model can increase memory usage and latency.
| Method | Wall-clock (s) | TFLOPs | Acc. | Acc. / TFLOPs |
|---|---|---|---|---|
| Cool-Fusion | 39.4 | 86 | 21.8 | 0.25 |
| AdaFuse | 33.9 | 74 | 43.7 | 0.60 |
| ThinkFuse | 68.7 | 77 | 72.9 | 0.95 |
Scroll sideways to see every column.
Efficiency comparison of ThinkFuse and comparable test-time fusion methods. Wall-clock time and TFLOPs are averaged per example, and accuracy is averaged over the five benchmarks, with the Qwen3-4B × Ministral pair.
Ablation Study
All ablations pair Qwen3-4B as the primary model with Ministral-3B-Reasoning as the auxiliary model. Without the budget penalty (\(\lambda = 0\)), accuracy falls to 88.5% on MATH-500 and 90.5% on GSM8K, below the standalone primary model, which indicates that the soft budget stabilizes ThinkFuse by discouraging excessive fusion interventions.
The default segment length \(T = 8\) yields the best results. A segment of \(T = 4\) fails to encapsulate a meaningful reasoning step, and longer segments (\(T = 16\) and \(T = 32\)) diminish the gains, which suggests that excessively long segments delay necessary interventions.
Scoring candidates by the primary model's PPL alone works best, while the average PPL drops MATH-500 to 80.5% and limits GSM8K to the standalone baseline (92.0%). The confidence gap reaches 98.6% on GSM8K but falls below the standalone baseline on MATH-500 (87.5%), and entropy, which preserves the gains on both datasets, is the default uncertainty signal.
| Setting | MATH-500 | GSM8K |
|---|---|---|
| Fusion budget penalty | ||
| Soft budget (ours, \(\lambda = 50\)) | 90.0(+1.0) | 96.9(+4.9) |
| No budget penalty (\(\lambda = 0\)) | 88.5(−0.5) | 90.5(−1.5) |
| Segment length \(T\) | ||
| \(T = 4\) | 87.0(−2.0) | 93.5(+1.5) |
| \(T = 8\) (ours) | 90.0(+1.0) | 96.9(+4.9) |
| \(T = 16\) | 88.0(−1.0) | 93.5(+1.5) |
| \(T = 32\) | 88.5(−0.5) | 93.5(+1.5) |
| Trajectory compatibility scoring | ||
| Primary PPL (ours) | 90.0(+1.0) | 96.9(+4.9) |
| Average PPL | 80.5(−8.5) | 92.0(+0.0) |
| Token-level uncertainty estimation | ||
| Entropy (ours) | 90.0(+1.0) | 96.9(+4.9) |
| Confidence gap | 87.5(−1.5) | 98.6(+6.6) |
Scroll sideways to see every column.
Impact of individual ThinkFuse components on downstream performance. Values in parentheses denote the accuracy delta relative to the standalone baseline. All experiments are conducted with Qwen3-4B as the primary model and Ministral-3B-Reasoning as the auxiliary model.
Case Study
On an AIME 2024 example, both traces share the initial coordinate setup. The standalone baseline prematurely grounds the coordinates of \(D\), failing to satisfy the second angle-bisector condition and trapping the reasoning in a dead-end quadratic, resulting in an incorrect answer. In contrast, ThinkFuse triggers an intervention and selects an auxiliary segment that keeps \(D\) and \(C\) symbolic and successfully enforces both distance-to-line constraints before solving.
Shared prefix. For AIME-2024 #17, both traces begin with the same coordinate setup: \(A=(0,0)\), \(B=(2,0)\), \(D=(p,q)\), \(C=(r,s)\), and \(M=((2+r)/2,s/2)\) is the midpoint of \(BC\). Both traces then try to use the two angle-bisector conditions at \(M\).
Generated trace near the divergence
⋯ I place the quadrilateral in coordinates and write the distance equations. Since \(M\) is the midpoint of \(BC\), it should lie on the two angle bisectors. ⋯
the valid solution is b = -7/sqrt(2), a = 7/sqrt(2)
So, D is at ((7*sqrt(2))/2, -7*sqrt(2)/2).
⋯ I plug this coordinate for \(D\) back into the remaining constraints and solve for \(C\). ⋯
(2+sqrt(2))*y^2 + 9*(sqrt(2)+1)*y
+ (22+7*sqrt(2)) = 0
// Comment: this fixes D as if ∠DAB = 45°. The second angle-bisector condition has not been used, so the trace enters a dead-end quadratic.
Answer: \(\boxed{0}\) ×
Generated trace near the fusion point
⋯ I keep \(D=(p,q)\) and \(C=(r,s)\) symbolic, then express the two angle-bisector conditions as distance-to-line equations. ⋯
[Two angle-bisector distance equations from M = (B+C)/2]
-8 - 2p + 3r = 0
=> r = (8 + 2p)/3
=> s = 2q/3
Substituting these relations keeps both angle-bisector constraints active while the remaining equations are solved.
// Comment: the auxiliary segment keeps D = (p, q) symbolic. It uses both angle-bisector conditions before solving.
Answer: \(\boxed{180}\) ✓
Qualitative comparison on AIME 2024. Standalone denotes Qwen3-4B. ThinkFuse uses Qwen3-4B as the primary model and Ministral-3B-Reasoning as the auxiliary model. Both traces start from the same coordinate formulation; the highlighted text marks the standalone step that enters a dead-end quadratic and the auxiliary symbolic step selected by ThinkFuse.
BibTeX
@article{kang2026thinkfuse,
title={ThinkFuse: Trajectory-Aware Test-Time Fusion for Small Reasoning Models},
author={Kang, Myunghoon and Lee, Jungseob and Seo, Jaehyung and Lim, Heuiseok},
journal={arXiv preprint arXiv:2610.07803},
year={2026}
}