Jungseob Lee Publications

Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models

arXiv preprint

Jungseob Lee1, Seongtae Hong1, Dongyub Jude Lee2, Chanjun Park3, Jaehyung Seo4, Sugyeong Eo5*, Heuiseok Lim1*

1Korea University 2Zoom Communications 3Soongsil University 4Konkuk University 5Yonsei University

*Corresponding authors

GLANCE drafts a whole block in one forward pass over the target's already fused vision-language states, and the target commits exactly its greedy output.

Two-part diagram of GLANCE. In the top part, a budget request summary page with row 4, Other Expense, outlined passes through a vision encoder into the frozen VLM target, together with the prompt, total amount of other expenses?. A dashed path that would carry the page's raw visual tokens, drawn as eight patches of the page, to the trained block head is crossed out. The target is drawn as a row of decoder layers, five of which are highlighted, and their outputs form a stack of fused states that feeds the block head. The bottom part shows one decoding round in three numbered steps. In 1, block draft, the head's input row holds the root b followed by masked positions, and one pass fills them with the tokens The, total, amount, of, other, each with a small bar chart of scores. In 2, candidate tree, a tree of budget N grows from b into The and A, with children total and sum under The and total and value under A. In 3, verify and commit, one target pass with an ancestor mask accepts the path b, The, total and crosses out the other nodes, and the round commits The and total followed by a new root, which equals the target greedy output. An arrow labeled next round returns to the block draft.
Overview of GLANCE. The block head (top) reads the target's fused vision-language states, never raw visual tokens. In one round (bottom), one draft pass fills the block, the top prefixes form a candidate tree, and one target pass verifies it.
Bar chart from Our World in Data. Share that agrees that vaccines are important for children to have, 2018: United Arab Emirates 94%, Mauritania 91%, Spain 88%, Armenia 73%, South Korea 72%.

PromptRead the chart and answer the question. State the exact values and labels you read off the chart, then explain. Question: How many colors are used in the graph?

  • GLANCE (1 draft pass) 3.1×over AR 7.74 ms/token4.9 tokens/round
    256 / 256 tokens1.97 s
  • EAGLE3-VL (8 draft passes) 2.4×over AR 10.29 ms/token3.8 tokens/round
    193 / 256 tokens1.97 s
  • AR autoregressive 1×baseline 24.3 ms/token1 token/pass
    82 / 256 tokens1.97 s

Qwen3-VL-8B on one RTX A6000

Abstract

Speculative decoding accelerates generation without changing its output, but on vision-language models (VLMs) a self-reinforcing cycle holds it back. Because an autoregressive drafter pays a sequential pass for each drafted token, it must stay small and can ill afford to attend to the image at each pass. Prior work therefore compresses or hides the image, leaving the drafter weakest on the text the image determines. We present GLANCE, a one-pass block drafter that breaks this cycle on an unmodified VLM target. Its block-diffusion head drafts a whole block in one forward pass over the target's already fused vision-language states, reading the multimodal context once, however deep the draft. The target verifies a wide candidate tree in one pass and commits exactly its greedy output. In one production engine at a fixed round budget, GLANCE decodes up to 3.05 times faster than autoregressive decoding and outpaces the production EAGLE3-VL head on average and by about 11% on grounded tasks. An entropy law explains when drafting pays, predicting the longest accepted blocks on grounded tasks, where the target's next-token entropy is lowest. Our code is available at https://github.com/js-lee-AI/GLANCE.

Method

In one decoding round of GLANCE, a block head drafts a whole block of future tokens in one forward pass, and the target verifies a wide tree of those candidates in one forward pass and commits its own greedy prefix. Only the head is trained, and the target stays frozen.

  1. Draft pass

    One-pass drafting on fused states

    GLANCE drafts with a block-diffusion head of 5 layers, 1.05B parameters, and block size \(B = 16\), attached to a frozen Qwen3-VL-8B target. The first block position holds the pending root \(b\) and the other \(B - 1\) positions are masked, so one forward pass conditioned on the cached prefix \(x\) returns a marginal \(q_j(\cdot \mid x, b)\) over the vocabulary at every offset \(j = 1, \dots, L\) with \(L = B - 1 = 15\). A candidate prefix is ranked by its plug-in score

    \[\hat\pi(y_{1:\ell}) = \prod_{j \le \ell} q_j(y_j \mid x, b)\]

    The head never reads raw visual tokens. Its cross-attention keys and values are the target's hidden states at five layers spread over the depth of the stack, \(\{1, 9, 17, 25, 33\}\). By the time these states exist, the target has already merged the image into its text representation, so the drafter inherits visual grounding without an encoder, an adaptor, or compressed image tokens of its own.

  2. Verify pass

    Wide-tree verification

    A candidate tree \(S\) is a prefix-closed set of nonempty strings after the root \(b\), of depth at most \(L\) and budget \(N = |S|\). The tree builder keeps the \(N\) prefixes with the highest plug-in score, and the tree is packed into one ancestor-masked target pass, in which the row of a depth-\(d\) node \(y_{1:d}\) yields exactly the target's distribution \(p(\cdot \mid x \circ b \circ y_{1:d})\). With \(Y\) the target's greedy continuation, the accepted length is

    \[A(S) = \max\{k \ge 0 : Y_{1:j} \in S \ \text{for all } j \le k\}\]

    The walk commits \(Y_{1:A(S)}\) and emits the next target-greedy token from the last verified row as the new root. A weak head can only shorten the accepted prefixes, and no fallback path is needed, because every committed token is the argmax of a target row. Growing \(N\) adds verifier work inside that single pass but no draft passes. The experiments use \(N = 63\), or a 31-node tree that together with the root matches the production head's budget of 32 draft tokens a round.

Assume the target's argmax is unique at every visited prefix and its top-2 logit gap there exceeds the numerical difference between packed and unpacked evaluation. Then, for any candidate tree, the walk commits exactly the target's greedy autoregressive sequence, token for token (Theorem 1 of the paper, stated informally).

In fp32, GLANCE is bitwise identical to greedy decoding on all 60 audited prompts. In bf16, the target does not reproduce its own greedy output across two runs even without any drafter, because kernel choice perturbs near-ties, and GLANCE reproduces that output at least as often as a second run of the target does. The exactness guarantee concerns greedy decoding. Under sampling, every committed token is a draw from the target's own distribution, and the output distribution is preserved, while bitwise identity no longer applies.

Only the block head is trained, while the vision tower and the target stay frozen. The training rows are the target's own greedy generations on 8,000 COCO-Caption2017 and 8,000 TextVQA prompts. The head is initialized from a released text-only block head for Qwen3-8B and trained for one epoch on one GPU. The recipe contains no document, infographic, or chart data.

An Entropy Law of Draftability

The paper characterizes how draftable a workload is with one scalar, the target's next-token entropy \(H\) at the round root. The matches along a block are modeled as a survival process with a single match probability \(p_{\mathrm{m}}(H)\) that varies slowly with the offset. Taking the match logit to be affine in the entropy gives

\[\begin{aligned}\operatorname{logit} p_{\mathrm{m}}(H) &= b_0 - b_1 H,\\ \mathbb{E}[a \mid H] &\approx \sum_{\ell=1}^{L} p_{\mathrm{m}}^{\ell} = \frac{p_{\mathrm{m}}\,(1 - p_{\mathrm{m}}^{L})}{1 - p_{\mathrm{m}}}\end{aligned}\]

with block horizon \(L = 15\), where \(a\) is the accepted length. Here \(b_0\) is the zero-entropy intercept and \(b_1 > 0\) the entropy slope, both fitted on the accepted lengths of each task's rounds with a 31-node tree.

Two predictions follow. First, \(\mathbb{E}[a \mid H]\) strictly decreases in \(H\), hence any workload that systematically lowers the target's entropy yields longer accepted blocks. Grounded generation is such a workload, because it copies determinate strings such as glyphs, numbers, and answer spans from the image. Second, the law has a ceiling that grounded generation breaks.

Two examples. Under the heading Document reading (low entropy), a budget request summary page with row 4, Other Expense, $975.00, outlined, the question total amount of other expenses?, and the drafted tokens The total amount of other expenses is $975.00, all marked committed, with an arrow from the outlined row to $975.00. Under the heading Open captioning (high entropy), a photograph of a seaplane on calm water in front of hills, the question describe the image., the committed tokens A seaplane floats, and three possible continuations, on calm water, near soft hills, and in a quiet bay.
Draftability of grounded and open generation. On a document (top), the image pins the answer and one draft pass commits the entire block. Open captioning (bottom) admits many continuations, and the same pass commits only a short prefix.

The law caps the near-certain mean at \(e^{b_0}\). If near-zero-entropy rounds enter, with probability \(\pi\), a verbatim-copy state whose matches persist, the mean as \(H \to 0\) becomes \(\pi L + (1 - \pi) M\), where \(M\) is the mean of the ordinary state, and it exceeds \(e^{b_0}\) once \(\pi\) passes an explicit threshold (Theorem 2 of the paper, stated informally). The copy regime is where one-pass drafting gains the most, because a verbatim run of length \(\ell\) costs an autoregressive drafter \(\ell\) sequential passes and a block drafter one.

Head-to-Head in a Production Engine

The target is Qwen3-VL-8B-Instruct, kept frozen. Five tasks are ordered from open-ended to grounded, namely COCO captioning, TextVQA, InfographicVQA (InfoVQA), DocVQA, and ChartQA. The last three, whose answers are read off a document or chart, are the grounded tasks, and they are the three tasks of lowest mean target entropy. Decoding is batch one and greedy, with up to 256 new tokens. Results are reported as the mean acceptance length \(\tau = \mathbb{E}[a] + 1\), so autoregressive (AR) decoding has \(\tau = 1\), and as wall-clock speedup, the ratio of the mean decode time for a token, autoregressive over speculative, with both arms measured in one engine on one GPU.

The table shows GLANCE and the production EAGLE3-VL head inside SGLang 0.5.6 on one RTX A6000 in bf16, with both CUDA-graph captured and both verifying a tree of 32 draft tokens a round, EAGLE3-VL its own top-8 dynamic tree and GLANCE its prefix tree. Engine, GPU, and round budget are thus shared, and the draft tokens come from eight sequential passes in EAGLE3-VL and from one pass in GLANCE. This eight-pass tree decodes 18 to 33% faster than the released configuration of EAGLE3-VL, three passes over four draft tokens, on every task.

Task\(\bar{H}\)AR
ms/tok
EAGLE3-VL, eight draft passesGLANCE, one draft passGLANCE
faster by
\(\tau\) ↑ms/tok ↓Speedup ↑\(\tau\) ↑ms/tok ↓Speedup ↑
Higher-entropy tasks
Captioning0.4724.723.5111.542.14×2.9213.101.89×−11.9%
TextVQA0.3924.344.319.562.55×3.5710.762.26×−11.2%
Lower-entropy tasks
InfographicVQA0.3124.583.5111.992.05×3.6510.842.27×+10.6%
DocVQA0.1524.913.6811.672.14×3.9110.512.37×+11.0%
ChartQA0.1924.184.508.792.75×4.697.933.05×+10.8%
Geometric mean2.31×2.34×+1.3%

Scroll sideways to see every column.

GLANCE against the production EAGLE3-VL head in SGLang, sharing engine, GPU, and a tree of 32 draft tokens, filled in eight passes by EAGLE3-VL and one by GLANCE. \(\bar{H}\) is the target's mean root entropy in nats. Blue marks GLANCE.

On the three lower-entropy tasks, GLANCE is the faster system by 10.6 to 11.0% and reaches 3.05× the speed of autoregressive decoding on ChartQA. Every paired bootstrap interval excludes zero, GLANCE is faster on at least 81 of the 101 prompts of each of these tasks, and it also leads on the five-task geometric mean, by 1.3% with a 95% interval from 0.3 to 2.2%. None of the three tasks appears in GLANCE's training data.

On captioning and TextVQA, the two tasks with the highest mean entropy \(\bar{H}\), the eight-pass head leads. The product-of-marginals ranking of the draft pass accounts for this split. It is accurate when the tokens of a block are nearly determined by the image, whereas on free-running text an autoregressive drafter stays coherent by construction. With an EAGLE-3 head and a GLANCE head trained on one corpus and run under the same tree, GLANCE is faster on all five tasks, by 4.5% on captioning, 14.9% on TextVQA, and 16.1 to 25.7% on the three grounded tasks.

The measurements are at batch one, the regime in which decoding is bound by memory bandwidth and speculative decoding is deployed for latency. Serving at large batch sizes, which changes the cost of verification, is not measured.

Acceptance Against Released Drafters

The table reports the acceptance length of every drafter released for these targets, each at its own operating point. Among trained heads, GLANCE accepts the longest blocks on all five tasks. Classic two-model speculation with a 4B draft accepts more on four of the five tasks, but a draft half the size of the target costs about half a target pass for each drafted token, so it slows decoding on every task.

Method (draft passes a round)Params
Acceptance length \(\tau\) ↑
Lossless
CaptioningTextVQAInfoVQADocVQAChartQA
Training-free and two-model drafting
n-gram lookup (0)01.292.492.533.302.57✓
Classic SD, Qwen3-VL-4B (8)4.4B3.533.793.954.394.47✓
Classic SD, Qwen3-1.7B text-only (8)2.0B1.491.421.901.682.12✓
Trained draft heads
EAGLE3-VL (5)0.40B2.132.662.482.562.92✓
EAGLE-2, ViSpec codebase (3)0.23B2.412.452.382.542.89✓†
ViSpec, official recipe (3)0.31B2.452.432.452.462.95✓†
Medusa, ViSpec codebase (1)0.08B1.511.521.471.531.61✓†
GLANCE (1)1.05B3.093.463.753.765.12✓†
Matched training
EAGLE-3 head, depth-3 chain (3)0.40B1.621.591.551.581.87✓
GLANCE, budget-63 tree (1)1.05B4.053.973.904.137.44✓†
ViSpec's training corpus
GLANCE, 1 epoch (1)1.05B3.023.443.753.785.02✓†
GLANCE, 21 epochs (1)1.05B3.043.493.783.855.07✓†
ViSpec's home target Qwen2.5-VL-7B
ViSpec, released head (3)0.35B3.343.263.213.043.59✓†
GLANCE (1)1.23B4.002.613.473.124.72✓†

Scroll sideways to see every column.

Acceptance length on Qwen3-VL-8B except the last group, each drafter at its own operating point, GLANCE with the budget-63 tree. Params counts drafter weights. ✓ marks exact acceptance, † output audited bitwise identical to greedy decoding. Blue marks GLANCE.

The two tables run the same GLANCE head, the production-engine table in SGLang with images at native resolution and a tree of 32 draft tokens, this one in the Hugging Face implementation at 896 pixels with the budget-63 tree, so their acceptance lengths differ.

The ViSpec codebase isolates what its vision adaptor adds. Trained without the adaptor, the head is exactly the EAGLE-2 drafter, and the two heads differ by at most 0.08 in acceptance length on any task, so the target's fused states, which both heads read, already carry what the adaptor's compressed image tokens add.

Four controls separate the architecture from its training. First, trained from scratch with one 26K-row corpus, target, global batch, schedule, and framework, GLANCE decodes faster than an EAGLE-3 head on all five tasks under the shared tree of the production-engine comparison. This corpus, unlike the primary recipe, contains document and chart rows. Second, retrained on ViSpec's own 68K-row corpus, at one epoch and at ViSpec's twenty-one, GLANCE stays within one percent of its acceptance under the primary recipe, pooled over the five tasks. Third, on ViSpec's home target Qwen2.5-VL-7B, a GLANCE head trained on that target accepts longer blocks than the released ViSpec head on four of the five tasks. Fourth, the released text head that GLANCE starts from, run unchanged, trails the production head on all five tasks of the production-engine table, by 15% on the geometric mean, and the training lengthens its accepted blocks by 17 to 21%, which turns that deficit into GLANCE's lead.

Where the Acceptance Comes From

GLANCE's head has 2.6× the parameters of the EAGLE3-VL head, so a larger head might explain the gains. Panel (a) measures what the tree adds by running the identical head as a width-1 chain. The budget-63 tree accepts between 1.45 and 1.49× the chain's length on every task, a nearly constant factor, as the law anticipates. The tree also decodes 1.36× faster than the chain in wall-clock time, net of verifying a wider tree. Under the shared budget of the production-engine table, a GLANCE round is 4 to 7% shorter than an EAGLE3-VL round on every task, so the larger head drafts all fifteen offsets in one pass for less than the small head pays for eight.

Three panels over the tasks Caption, TextVQA, InfoVQA, DocVQA, and ChartQA. Panel (a) is a bar chart of acceptance length for the identical head run as a width-1 chain and as a budget-63 tree. The chain bars rise from about 2.0 on Caption to about 3.5 on ChartQA, and the tree bars from about 3.0 to about 5.2, higher than the chain on every task. Panel (b) is a bar chart of acceptance length when the head reads text-only states or fused vision-language states. The fused bars are higher on every task, about 2.9 against 2.8 on Caption, about 3.6 against 2.9 on DocVQA, and about 4.8 against 4.0 on ChartQA. Panel (c) is a line plot of speedup over autoregressive decoding against the verifier budget N at 15, 31, 47, and 63. Every task rises from 15 to 31 and then flattens, with ChartQA highest, from about 3.1 to about 3.4, and Caption lowest, from about 1.9 to about 2.05.
Sources of GLANCE's acceptance. (a) The identical head as a width-1 chain and as a budget-63 tree. (b) The head reading the target's fused states or a text-only model's states, at budget 31. (c) Speedup against the verifier budget \(N\).

Panel (b) replaces the target's fused states with those of a text-only Qwen3-8B, which never sees the image and whose states the head was initialized on. Acceptance falls on every task, and it falls most on grounded ones, since the text-only states retain 97% of GLANCE's acceptance on captioning but only 80% on DocVQA. Zeroing the conditioning states collapses \(\tau\) to 1.11, which shows that the head drafts from the target's context.

Panel (c) sweeps the verifier budget. Speedup rises from \(N = 15\) to 31 on every task and then flattens. Width pays the most on ChartQA, the task with the longest accepted blocks.

Testing the Entropy Law

Panel (a) shows GLANCE's accepted length falling with entropy on all five tasks, with grounded tasks accepting more at nearly every entropy level. The first prediction of the law holds, since isotonic fits explain at least 96% of the variance of the ten entropy-decile means on every task. The fitted slope \(b_1\) steepens with grounding, from 0.46 on captioning to 1.71 on ChartQA.

Three panels. Panel (a) is a line plot of accepted length against next-token entropy in nats on a log scale, by entropy decile, for Caption, TextVQA, InfoVQA, DocVQA, and ChartQA. Every curve falls as entropy grows, from between about 2.5 and 7 near zero entropy to between about 1 and 2 at about 1 nat, and ChartQA is highest over most of the range. Panel (b) is a bar chart of the acceptance gain from the correct image in percent with 95% bootstrap intervals, about 3.5 on Caption, 6.8 on TextVQA, and 11.7 on DocVQA. Panel (c) is a bar chart of the law's cap against the measured floor of accepted length. The measured floor is higher on every task, about 2.5 against 2.3 on Caption, 3.2 against 2.5 on TextVQA, 4.4 against 3.2 on InfoVQA, 5.2 against 3.4 on DocVQA, and 6.9 against 4.5 on ChartQA.
Tests of the entropy law. (a) Accepted length by decile of the target's next-token entropy. (b) Acceptance gain from the correct over a mismatched image, with 95% bootstrap intervals. (c) The fitted law's cap \(e^{b_0}\) against the measured lowest-decile mean.

To test whether the image itself is responsible, the image alone is varied. The same prompts are decoded with the correct and with a mismatched image along a shared teacher-forced trajectory. Panel (b) shows that the correct image lengthens accepted blocks on every task, and more so as grounding increases. The correct image also lowers the target's mean entropy along that trajectory, from 0.72 to 0.26 nats on DocVQA, and once entropy is held fixed the remaining image effect is no longer positive. Together with the fused-state ablation, this places the image's contribution in the entropy channel.

Panel (c) compares the cap \(e^{b_0}\) with the measured mean of \(a\) in each task's lowest-entropy decile. Every floor lies above its cap, and the excess grows with grounding, reaching 5.23 against 3.41 on DocVQA. On ChartQA, 11.5% of all rounds accept eight or more tokens.

The law keeps its form under a change of drafter, target, and modality, with a positive slope in every case. Two-model drafts on Qwen2.5-VL-7B and Qwen2-VL-7B fit at \(b_1 = 0.67\) and 0.40, and beyond vision the slope stays positive on speech recognition, chat, code, and a non-Qwen backbone. The characterization of draftability is stated for targets that feed vision as a token prefix, which covers the Qwen-VL family studied here.

BibTeX

@misc{lee2026glance,
  title = {Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models},
  author = {Jungseob Lee and Seongtae Hong and Dongyub Jude Lee and Chanjun Park and Jaehyung Seo and Sugyeong Eo and Heuiseok Lim},
  year = {2026},
  journal = {arXiv preprint},
  eprint = {2609.00355},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.00355},
}