Jungseob Lee Publications

To Isolate or to Score? Model-Adaptive Assessment for Cost-Efficient Multi-Agent RAG

arXiv preprint

Jungseob Lee1, Chanjun Park2*, Heuiseok Lim1*

1Korea University 2Soongsil University

*Corresponding authors

A training-free pipeline that diagnoses each model–task pair with a one-time probe and routes weak baselines to per-document isolation and strong baselines to a scoring treatment.

Diagram in three phases. Phase 1, model profiling (offline): a target LLM and calibration data (n = 100) feed a one-time RSC probe and evaluation, which outputs the context capacity EM NF, the scoring trend rho star, and the coupling strength rho hat 1. Phase 2, MADARA router (decision logic): Path A, weak multi-doc capacity with EM NF below 30%, leads to PDE. Path B, strong multi-doc capacity with EM NF of at least 30%, splits into stochastic scoring with rho star above minus 1.0, which leads to SDA, and rho star equal to minus 1.0, where rho hat 1 below 0.5 (weakly coupled quality-ordered) leads to CoT and rho hat 1 of at least 0.5 (strongly coupled quality-ordered) leads to ATF. Phase 3, treatment execution (online processing): PDE, per-document extraction, passes documents D1 to D5 separately to the model and combines the five answers by voting. SDA, score distribution alignment, turns score histograms into percentile ranks. CoT, de-polarization, places thinking and reasoning between the prompt and the score. ATF, adaptive threshold filtering, keeps the documents above the threshold mu minus 0.5 sigma for the generator.
The MADARA Model-Adaptive Routing Architecture. (Left) A one-time RSC probe evaluates the target LLM's context capacity (\(\text{EM}_{\text{NF}}\)) and scoring behavior (\(\rho^*, \hat{\rho}_1\)). (Middle) The router identifies a capability phase-transition: weak-baseline models are strictly routed to bypass multi-document evaluation. (Right) PDE (Isolation) structurally separates documents to cure context confusion, while scoring-only treatments (SDA, CoT, ATF) refine document assessment for strong models.

Abstract

Multi-agent document assessment for retrieval-augmented generation is computationally expensive, driving practitioners toward smaller, deployable models whose assessment mechanisms remain poorly understood. We conduct a controlled study of training-free interventions on 7B–9B instruction-tuned models across diverse QA benchmarks, revealing a sharp dichotomy in how models benefit from assessment. For weaker baselines, the dominant mechanism is per-document isolation. Astoundingly, assessment-free isolation matches full multi-agent assessment, demonstrating that resolving multi-document context confusion, rather than scoring quality, drives outsized gains of up to 50 percentage points. Conversely, for strong baselines where scoring quality matters, we introduce Reasoning-Score Coupling, a label-free perturbation probe that classifies scoring behavior. Integrating these findings, we propose MADARA, a model-adaptive routing architecture. Crucially, MADARA's diagnostic thresholds derived from a single pilot model generalize zero-shot to four unseen model families, providing a robust, lightweight pipeline to eliminate computational overhead.

Method

MADARA is a Diagnose→Treat pipeline over a three-agent document assessment framework. It diagnoses each model–task pair on 100 calibration queries by its accuracy under the No-Filter (NF) baseline, which passes all documents to the generator, and by its Reasoning-Score Coupling (RSC).

RSC compares a model's document scores under chain-of-thought (CoT) reasoning with those obtained after three levels of increasing perturbation (shuffled, contradicted, and random reasoning). With \(\hat{\rho}_k\) the Spearman rank correlation between normal and perturbed scores at level \(k\), the trend coefficient is

\[\rho^* = \text{Spearman}\bigl([1,\,2,\,3],\;[\hat{\rho}_1,\,\hat{\rho}_2,\,\hat{\rho}_3]\bigr)\]

A pair is quality-ordered if \(\rho^* = -1.0\), which requires \(\hat{\rho}_1 > \hat{\rho}_2 > \hat{\rho}_3\), and stochastic otherwise. The protocol requires no gold labels, and \(\hat{\rho}_1\) measures the strength of baseline coupling.

  1. Path A (weak multi-doc capacity)

    Per-Document Extraction (PDE)

    If \(\text{EM}_{\text{NF}} < \tau_{\text{NF}}\), the model struggles with multi-document context. After three-agent scoring, it generates an answer from each top-\(k\) document individually. Candidates are grouped by normalized string match, and the group with the highest cumulative score is selected.

  2. Path B (stochastic scoring)

    Score Distribution Alignment (SDA)

    If \(\rho^* > -1.0\), reasoning quality does not drive scores. SDA bypasses reasoning entirely. It converts each agent's raw scores to percentile ranks, then aggregates them via weighted averaging to produce a calibrated ranking.

  3. Path B (weakly coupled quality-ordered)

    CoT De-Polarization

    If \(\rho^* = -1.0\) and \(\hat{\rho}_1 < 0.5\), the baseline failure mode is polarization, in which agents without reasoning guidance assign extreme scores (0 or 5) to about 80% of documents. Agents therefore generate explicit reasoning before scoring.

  4. Path B (strongly coupled quality-ordered)

    Adaptive Threshold Filtering (ATF)

    If \(\rho^* = -1.0\) and \(\hat{\rho}_1 \geq 0.5\), baselines already produce effective rankings. ATF filters rather than reranks. With \(\tau = \mu(\mathbf{s}) - 0.5\,\sigma(\mathbf{s})\), it retains only documents with \(s_i \geq \tau\) and requires no additional LLM calls.

The routing thresholds (\(\rho^* = -1.0\), \(\hat{\rho}_1 = 0.5\), \(\tau_{\text{NF}} = 30\%\)) were derived from a single pilot model, Mistral-7B, and applied zero-shot to all others.

Main Results

Five open-weight, instruction-tuned 7B–9B models are evaluated on sampled subsets (∼1K queries) of CONFLICTS, an adversarial QA benchmark with inter-document contradictions, and FEVER, a binary fact-verification task.

TaskDiagnosticsComponent Methods EM (%)MADARA (Ours)
RSC\(\hat{\rho}_1\)NF3-AgentCoTSDAStrategyEM\(\Delta_{\text{NF}}\)
Llama-3.1-8B
ConflictsQuality-Ordered0.7613.924.116.519.4PDE50.2***+36.3
FEVERQuality-Ordered0.6588.787.888.488.2ATF90.9*+2.2
Mistral-7B-v0.3
ConflictsQuality-Ordered0.3518.118.622.820.3PDE43.5***+25.4
FEVERQuality-Ordered0.2890.790.992.690.7CoT92.6+1.9
Qwen3-8B
ConflictsStochastic0.3960.862.462.463.0SDA63.0+2.2
FEVERQuality-Ordered0.6389.689.889.590.0ATF90.9+1.3
Qwen2.5-7B
ConflictsStochastic0.4659.965.062.465.4SDA65.4+5.5
FEVERQuality-Ordered0.6287.188.987.590.5**ATF†88.2+1.1
Gemma-2-9B
ConflictsStochastic0.7960.864.663.363.3SDA63.3+2.5
FEVERQuality-Ordered0.5892.292.492.592.5ATF93.6+1.4
Average
Conflicts42.746.945.546.3–57.1+14.4
FEVER89.790.090.190.4–91.2+1.5

Scroll sideways to see every column.

MADARA dynamically routes models to optimal assessment strategies, maximizing Exact Match (EM%). The Strategy is determined zero-shot via RSC and No-Filter (NF) baselines. Significance vs. NF (McNemar's test with Holm-Bonferroni): \(^{*}p{<}0.05\), \(^{**}p{<}0.01\), \(^{***}p{<}0.001\). †Routed to ATF due to crossing the baseline coupling threshold (\(\hat{\rho}_1 = 0.62 \ge 0.5\)), narrowly missing the optimal SDA.

Routing sends the two weak-baseline pairs to PDE, which yields +36.3pp for Llama and +25.4pp for Mistral on CONFLICTS (\(p < 0.001\)). Averaged over the five models, the routed strategy reaches 57.1% EM on CONFLICTS against 42.7% for NF, and 91.2% against 89.7% on FEVER.

Isolation vs. Scoring

PDE-Random selects documents at random and votes uniformly, which bypasses multi-agent assessment. For Llama it matches the full PDE pipeline on both CONFLICTS (50.6% vs. 50.2%) and the held-out TriviaQA (79.6% vs. 79.6%). For Mistral, assessment-guided selection adds +19pp beyond random isolation on CONFLICTS.

Strong-baseline models (\(\text{EM}_{\text{NF}} \geq 60\%\)) show no benefit from PDE. Qwen3 on CONFLICTS, for example, drops from 60.8% to 58.6%.

TaskBasePDE Component Ablation
NF(1× cost)Rand.(1× cost)Unif.(≈4×)Full(≈4×)
Llama-3.1
Conflicts13.950.651.150.2
TriviaQA29.879.680.179.6
Mistral-v0.3
Conflicts18.124.538.843.5
TriviaQA67.671.573.773.7

Scroll sideways to see every column.

PDE component ablation (EM%) reveals that multi-agent assessment is redundant for weak baselines. For Llama, completely assessment-free isolation (Rand.) yields identical outsized gains as the computationally heavy Full pipeline, proving that resolving context confusion drives the improvement. (Cost multipliers indicate relative inference calls vs. standard RAG. CFL=Conflicts, TQA=TriviaQA; NF=No Filter; Unif.=assessment-guided + uniform vote; Full=score-weighted vote.)

RSC Diagnostic Results

Strong-baseline models lose quality-ordered scoring on the adversarial and factoid-QA tasks but maintain it on the simpler binary FEVER task. On the held-out TriviaQA, RSC classifications replicate those on CONFLICTS, except for Llama, whose scoring is degenerate.

TaskCorrelation (\(\hat{\rho}_k\))Trend (\(\rho^*\))RSC Class
ShuffledContra.Random
Llama-3.1-8B
Conflicts.76.43.08−1.0Quality-Ordered
FEVER.65.49.10−1.0Quality-Ordered
TriviaQADegenerate scoring (mean=0.17/5.0)‡
Mistral-7B-v0.3
Conflicts.35.22.18−1.0Quality-Ordered
FEVER.28.09.01−1.0Quality-Ordered
TriviaQA.47.05.02−1.0Quality-Ordered
Qwen3-8B
Conflicts.39.04.14−0.5Stochastic
FEVER.63.51.14−1.0Quality-Ordered
TriviaQA.64.45.45−0.5Stochastic
Qwen2.5-7B
Conflicts.46−.26.08−0.5Stochastic
FEVER.62.21.06−1.0Quality-Ordered
TriviaQA.55−.28.15−0.5Stochastic
Gemma-2-9B
Conflicts.79.40.40−0.5Stochastic†
FEVER.58.41.18−1.0Quality-Ordered
TriviaQA.83.19.48−0.5Stochastic

Scroll sideways to see every column.

RSC diagnostic results reveal that scoring behavior is a model-task interaction. Spearman correlations (\(\hat{\rho}_k\)) are shown under three increasing perturbation levels. Perfect monotonic degradation yields a trend coefficient of \(\rho^* = -1.0\), classifying the model-task pair as Quality-Ordered; otherwise, it is Stochastic. †Aggregate vs. per-query disagreement. ‡Scores degenerate at the absolute floor, making perturbation uninformative.

Computed independently for each of the 100 probe queries on CONFLICTS, Llama's mean correlations are strictly ordered (0.72 > 0.38 > −0.01), while Qwen3's means violate monotonicity.

Three violin plots of per-query Spearman rho, for Mistral-7B, Llama-3.1-8B, and Qwen3-8B, each with one violin for the shuffled, the contradicted, and the random perturbation level. Mistral-7B has means of plus 0.29, plus 0.18, and plus 0.08, with a green box reading trend rho of minus 1.00, p below 0.001. Llama-3.1-8B has means of plus 0.72, plus 0.38, and minus 0.01, with a green box reading trend rho of minus 1.00, p below 0.001. Qwen3-8B has means of plus 0.30, plus 0.03, and plus 0.16, with a red box reading trend rho of minus 0.50, p of 0.667.
Per-query Spearman \(\rho\) distributions on CONFLICTS. Violin widths show density, and horizontal lines mark medians. Numbers below violins indicate mean \(\rho\) values. Green and red boxes denote significant (\(\rho^* = -1.0\)) and non-significant monotonic trends, respectively.

Robustness Across Retrieval Quality

On TriviaQA, upgrading BM25 to a dense retriever (Contriever, 500 queries) shrinks Llama's PDE gain from +49.8pp to +20.8pp, and under a generative reranker (Qwen3-0.6B) the gain is +50.4pp.

Under dense retrieval, Llama's NF baseline (34.6%) marginally crosses \(\tau_{\text{NF}} = 30\%\), so a strict application of the router would bypass PDE and miss the +20.8pp isolation gain.

ModelPDE Gain (Δ EM vs. NF Baseline)
BM25(Sparse)Contriever(Dense)Qwen3-0.6B(Reranker)
Weak-Baseline Capacity
Llama-3.1-8B+49.8+20.8+50.4
Strong-Baseline Capacity
Qwen2.5-7B−0.5+3.6−1.0
Gemma-2-9B−0.8+2.4+0.8

Scroll sideways to see every column.

Impact of Context Quality Upgrades on TriviaQA. Upgrading retrieval quality (Dense/Reranker) reduces the isolation benefit for weak models, yet PDE remains mandatory to cure severe context confusion (e.g., \(+50.4\text{pp}\) for Llama). Conversely, strong models (Qwen, Gemma) possess intrinsic context capacity, rendering PDE unnecessary.

Multi-Hop Generalization

On MuSiQue, scoring treatments become optimal while isolation alone fails. For Qwen2.5-7B on 100 queries with 20 mixed supporting and distractor paragraphs per query, 3-Agent and SDA both improve EM by +16.0pp, while PDE-Random falls below NF. PDE lags the best scoring treatments by 6pp because isolated per-document voting cannot reconstruct missing chain steps.

The static threshold would still route this pair to PDE, a limitation of the current routing implementation.

MethodEM (%)F1 (%)Δ EM
NF (No-Filter)14.021.3—
3-Agent Baseline30.039.6+16.0
CoT De-Polarization21.030.9+7.0
SDA (rank)30.038.7+16.0
PDE24.034.4+10.0
PDE-Random12.019.6−2.0

Scroll sideways to see every column.

MuSiQue results (Qwen2.5-7B-Instruct, 100 queries). Bootstrap CI for the \(+16.0\)pp 3-Agent gain over NF: 95% CI \([+5.0, +27.0]\)pp, \(p=0.003\) (10,000 resamples).

BibTeX

@misc{lee2026madara,
  title = {To Isolate or to Score? Model-Adaptive Assessment for Cost-Efficient Multi-Agent RAG},
  author = {Jungseob Lee and Chanjun Park and Heuiseok Lim},
  year = {2026},
  journal = {arXiv preprint},
  eprint = {2606.25191},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.25191},
}