To Isolate or to Score? Model-Adaptive Assessment for Cost-Efficient Multi-Agent RAG
arXiv preprint
1Korea University 2Soongsil University
*Corresponding authors
Abstract
Multi-agent document assessment for retrieval-augmented generation is computationally expensive, driving practitioners toward smaller, deployable models whose assessment mechanisms remain poorly understood. We conduct a controlled study of training-free interventions on 7B–9B instruction-tuned models across diverse QA benchmarks, revealing a sharp dichotomy in how models benefit from assessment. For weaker baselines, the dominant mechanism is per-document isolation. Astoundingly, assessment-free isolation matches full multi-agent assessment, demonstrating that resolving multi-document context confusion, rather than scoring quality, drives outsized gains of up to 50 percentage points. Conversely, for strong baselines where scoring quality matters, we introduce Reasoning-Score Coupling, a label-free perturbation probe that classifies scoring behavior. Integrating these findings, we propose MADARA, a model-adaptive routing architecture. Crucially, MADARA's diagnostic thresholds derived from a single pilot model generalize zero-shot to four unseen model families, providing a robust, lightweight pipeline to eliminate computational overhead.
Method
MADARA is a Diagnose→Treat pipeline over a three-agent document assessment framework. It diagnoses each model–task pair on 100 calibration queries by its accuracy under the No-Filter (NF) baseline, which passes all documents to the generator, and by its Reasoning-Score Coupling (RSC).
RSC compares a model's document scores under chain-of-thought (CoT) reasoning with those obtained after three levels of increasing perturbation (shuffled, contradicted, and random reasoning). With \(\hat{\rho}_k\) the Spearman rank correlation between normal and perturbed scores at level \(k\), the trend coefficient is
\[\rho^* = \text{Spearman}\bigl([1,\,2,\,3],\;[\hat{\rho}_1,\,\hat{\rho}_2,\,\hat{\rho}_3]\bigr)\]
A pair is quality-ordered if \(\rho^* = -1.0\), which requires \(\hat{\rho}_1 > \hat{\rho}_2 > \hat{\rho}_3\), and stochastic otherwise. The protocol requires no gold labels, and \(\hat{\rho}_1\) measures the strength of baseline coupling.
-
Path A (weak multi-doc capacity)
Per-Document Extraction (PDE)
If \(\text{EM}_{\text{NF}} < \tau_{\text{NF}}\), the model struggles with multi-document context. After three-agent scoring, it generates an answer from each top-\(k\) document individually. Candidates are grouped by normalized string match, and the group with the highest cumulative score is selected.
-
Path B (stochastic scoring)
Score Distribution Alignment (SDA)
If \(\rho^* > -1.0\), reasoning quality does not drive scores. SDA bypasses reasoning entirely. It converts each agent's raw scores to percentile ranks, then aggregates them via weighted averaging to produce a calibrated ranking.
-
Path B (weakly coupled quality-ordered)
CoT De-Polarization
If \(\rho^* = -1.0\) and \(\hat{\rho}_1 < 0.5\), the baseline failure mode is polarization, in which agents without reasoning guidance assign extreme scores (0 or 5) to about 80% of documents. Agents therefore generate explicit reasoning before scoring.
-
Path B (strongly coupled quality-ordered)
Adaptive Threshold Filtering (ATF)
If \(\rho^* = -1.0\) and \(\hat{\rho}_1 \geq 0.5\), baselines already produce effective rankings. ATF filters rather than reranks. With \(\tau = \mu(\mathbf{s}) - 0.5\,\sigma(\mathbf{s})\), it retains only documents with \(s_i \geq \tau\) and requires no additional LLM calls.
The routing thresholds (\(\rho^* = -1.0\), \(\hat{\rho}_1 = 0.5\), \(\tau_{\text{NF}} = 30\%\)) were derived from a single pilot model, Mistral-7B, and applied zero-shot to all others.
Main Results
Five open-weight, instruction-tuned 7B–9B models are evaluated on sampled subsets (∼1K queries) of CONFLICTS, an adversarial QA benchmark with inter-document contradictions, and FEVER, a binary fact-verification task.
| Task | Diagnostics | Component Methods EM (%) | MADARA (Ours) | ||||||
|---|---|---|---|---|---|---|---|---|---|
| RSC | \(\hat{\rho}_1\) | NF | 3-Agent | CoT | SDA | Strategy | EM | \(\Delta_{\text{NF}}\) | |
| Llama-3.1-8B | |||||||||
| Conflicts | Quality-Ordered | 0.76 | 13.9 | 24.1 | 16.5 | 19.4 | PDE | 50.2*** | +36.3 |
| FEVER | Quality-Ordered | 0.65 | 88.7 | 87.8 | 88.4 | 88.2 | ATF | 90.9* | +2.2 |
| Mistral-7B-v0.3 | |||||||||
| Conflicts | Quality-Ordered | 0.35 | 18.1 | 18.6 | 22.8 | 20.3 | PDE | 43.5*** | +25.4 |
| FEVER | Quality-Ordered | 0.28 | 90.7 | 90.9 | 92.6 | 90.7 | CoT | 92.6 | +1.9 |
| Qwen3-8B | |||||||||
| Conflicts | Stochastic | 0.39 | 60.8 | 62.4 | 62.4 | 63.0 | SDA | 63.0 | +2.2 |
| FEVER | Quality-Ordered | 0.63 | 89.6 | 89.8 | 89.5 | 90.0 | ATF | 90.9 | +1.3 |
| Qwen2.5-7B | |||||||||
| Conflicts | Stochastic | 0.46 | 59.9 | 65.0 | 62.4 | 65.4 | SDA | 65.4 | +5.5 |
| FEVER | Quality-Ordered | 0.62 | 87.1 | 88.9 | 87.5 | 90.5** | ATF† | 88.2 | +1.1 |
| Gemma-2-9B | |||||||||
| Conflicts | Stochastic | 0.79 | 60.8 | 64.6 | 63.3 | 63.3 | SDA | 63.3 | +2.5 |
| FEVER | Quality-Ordered | 0.58 | 92.2 | 92.4 | 92.5 | 92.5 | ATF | 93.6 | +1.4 |
| Average | |||||||||
| Conflicts | 42.7 | 46.9 | 45.5 | 46.3 | – | 57.1 | +14.4 | ||
| FEVER | 89.7 | 90.0 | 90.1 | 90.4 | – | 91.2 | +1.5 | ||
Scroll sideways to see every column.
MADARA dynamically routes models to optimal assessment strategies, maximizing Exact Match (EM%). The Strategy is determined zero-shot via RSC and No-Filter (NF) baselines. Significance vs. NF (McNemar's test with Holm-Bonferroni): \(^{*}p{<}0.05\), \(^{**}p{<}0.01\), \(^{***}p{<}0.001\). †Routed to ATF due to crossing the baseline coupling threshold (\(\hat{\rho}_1 = 0.62 \ge 0.5\)), narrowly missing the optimal SDA.
Routing sends the two weak-baseline pairs to PDE, which yields +36.3pp for Llama and +25.4pp for Mistral on CONFLICTS (\(p < 0.001\)). Averaged over the five models, the routed strategy reaches 57.1% EM on CONFLICTS against 42.7% for NF, and 91.2% against 89.7% on FEVER.
Isolation vs. Scoring
PDE-Random selects documents at random and votes uniformly, which bypasses multi-agent assessment. For Llama it matches the full PDE pipeline on both CONFLICTS (50.6% vs. 50.2%) and the held-out TriviaQA (79.6% vs. 79.6%). For Mistral, assessment-guided selection adds +19pp beyond random isolation on CONFLICTS.
Strong-baseline models (\(\text{EM}_{\text{NF}} \geq 60\%\)) show no benefit from PDE. Qwen3 on CONFLICTS, for example, drops from 60.8% to 58.6%.
| Task | Base | PDE Component Ablation | ||
|---|---|---|---|---|
| NF(1× cost) | Rand.(1× cost) | Unif.(≈4×) | Full(≈4×) | |
| Llama-3.1 | ||||
| Conflicts | 13.9 | 50.6 | 51.1 | 50.2 |
| TriviaQA | 29.8 | 79.6 | 80.1 | 79.6 |
| Mistral-v0.3 | ||||
| Conflicts | 18.1 | 24.5 | 38.8 | 43.5 |
| TriviaQA | 67.6 | 71.5 | 73.7 | 73.7 |
Scroll sideways to see every column.
PDE component ablation (EM%) reveals that multi-agent assessment is redundant for weak baselines. For Llama, completely assessment-free isolation (Rand.) yields identical outsized gains as the computationally heavy Full pipeline, proving that resolving context confusion drives the improvement. (Cost multipliers indicate relative inference calls vs. standard RAG. CFL=Conflicts, TQA=TriviaQA; NF=No Filter; Unif.=assessment-guided + uniform vote; Full=score-weighted vote.)
RSC Diagnostic Results
Strong-baseline models lose quality-ordered scoring on the adversarial and factoid-QA tasks but maintain it on the simpler binary FEVER task. On the held-out TriviaQA, RSC classifications replicate those on CONFLICTS, except for Llama, whose scoring is degenerate.
| Task | Correlation (\(\hat{\rho}_k\)) | Trend (\(\rho^*\)) | RSC Class | ||
|---|---|---|---|---|---|
| Shuffled | Contra. | Random | |||
| Llama-3.1-8B | |||||
| Conflicts | .76 | .43 | .08 | −1.0 | Quality-Ordered |
| FEVER | .65 | .49 | .10 | −1.0 | Quality-Ordered |
| TriviaQA | Degenerate scoring (mean=0.17/5.0)‡ | ||||
| Mistral-7B-v0.3 | |||||
| Conflicts | .35 | .22 | .18 | −1.0 | Quality-Ordered |
| FEVER | .28 | .09 | .01 | −1.0 | Quality-Ordered |
| TriviaQA | .47 | .05 | .02 | −1.0 | Quality-Ordered |
| Qwen3-8B | |||||
| Conflicts | .39 | .04 | .14 | −0.5 | Stochastic |
| FEVER | .63 | .51 | .14 | −1.0 | Quality-Ordered |
| TriviaQA | .64 | .45 | .45 | −0.5 | Stochastic |
| Qwen2.5-7B | |||||
| Conflicts | .46 | −.26 | .08 | −0.5 | Stochastic |
| FEVER | .62 | .21 | .06 | −1.0 | Quality-Ordered |
| TriviaQA | .55 | −.28 | .15 | −0.5 | Stochastic |
| Gemma-2-9B | |||||
| Conflicts | .79 | .40 | .40 | −0.5 | Stochastic† |
| FEVER | .58 | .41 | .18 | −1.0 | Quality-Ordered |
| TriviaQA | .83 | .19 | .48 | −0.5 | Stochastic |
Scroll sideways to see every column.
RSC diagnostic results reveal that scoring behavior is a model-task interaction. Spearman correlations (\(\hat{\rho}_k\)) are shown under three increasing perturbation levels. Perfect monotonic degradation yields a trend coefficient of \(\rho^* = -1.0\), classifying the model-task pair as Quality-Ordered; otherwise, it is Stochastic. †Aggregate vs. per-query disagreement. ‡Scores degenerate at the absolute floor, making perturbation uninformative.
Computed independently for each of the 100 probe queries on CONFLICTS, Llama's mean correlations are strictly ordered (0.72 > 0.38 > −0.01), while Qwen3's means violate monotonicity.
Robustness Across Retrieval Quality
On TriviaQA, upgrading BM25 to a dense retriever (Contriever, 500 queries) shrinks Llama's PDE gain from +49.8pp to +20.8pp, and under a generative reranker (Qwen3-0.6B) the gain is +50.4pp.
Under dense retrieval, Llama's NF baseline (34.6%) marginally crosses \(\tau_{\text{NF}} = 30\%\), so a strict application of the router would bypass PDE and miss the +20.8pp isolation gain.
| Model | PDE Gain (Δ EM vs. NF Baseline) | ||
|---|---|---|---|
| BM25(Sparse) | Contriever(Dense) | Qwen3-0.6B(Reranker) | |
| Weak-Baseline Capacity | |||
| Llama-3.1-8B | +49.8 | +20.8 | +50.4 |
| Strong-Baseline Capacity | |||
| Qwen2.5-7B | −0.5 | +3.6 | −1.0 |
| Gemma-2-9B | −0.8 | +2.4 | +0.8 |
Scroll sideways to see every column.
Impact of Context Quality Upgrades on TriviaQA. Upgrading retrieval quality (Dense/Reranker) reduces the isolation benefit for weak models, yet PDE remains mandatory to cure severe context confusion (e.g., \(+50.4\text{pp}\) for Llama). Conversely, strong models (Qwen, Gemma) possess intrinsic context capacity, rendering PDE unnecessary.
Multi-Hop Generalization
On MuSiQue, scoring treatments become optimal while isolation alone fails. For Qwen2.5-7B on 100 queries with 20 mixed supporting and distractor paragraphs per query, 3-Agent and SDA both improve EM by +16.0pp, while PDE-Random falls below NF. PDE lags the best scoring treatments by 6pp because isolated per-document voting cannot reconstruct missing chain steps.
The static threshold would still route this pair to PDE, a limitation of the current routing implementation.
| Method | EM (%) | F1 (%) | Δ EM |
|---|---|---|---|
| NF (No-Filter) | 14.0 | 21.3 | — |
| 3-Agent Baseline | 30.0 | 39.6 | +16.0 |
| CoT De-Polarization | 21.0 | 30.9 | +7.0 |
| SDA (rank) | 30.0 | 38.7 | +16.0 |
| PDE | 24.0 | 34.4 | +10.0 |
| PDE-Random | 12.0 | 19.6 | −2.0 |
Scroll sideways to see every column.
MuSiQue results (Qwen2.5-7B-Instruct, 100 queries). Bootstrap CI for the \(+16.0\)pp 3-Agent gain over NF: 95% CI \([+5.0, +27.0]\)pp, \(p=0.003\) (10,000 resamples).
BibTeX
@misc{lee2026madara,
title = {To Isolate or to Score? Model-Adaptive Assessment for Cost-Efficient Multi-Agent RAG},
author = {Jungseob Lee and Chanjun Park and Heuiseok Lim},
year = {2026},
journal = {arXiv preprint},
eprint = {2606.25191},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.25191},
}