Jungseob Lee Publications

Where Activation Sparsity and KV-Cache Sparsity Cross in LLM Decoding

arXiv preprint

Jungseob Lee1, Seungyoon Lee1, Seongtae Hong1, Sugyeong Eo2†, Heuiseok Lim1†

1Korea University 2Yonsei University Mirae Campus

†Corresponding authors

A byte account of the decode step gives the context length at which the savings of activation sparsity and KV-cache sparsity are equal, from model dimensions and keep ratios alone.

Three decode steps drawn along a growing context, each with an attention block and a feed-forward block, ending in the next token. A projection-sparsity panel above them shows a row of activation channels over a grid of shared projection weights, where the weight columns selected at each step are filled blue and a blue arrow leads from them into the step. A KV-cache sparsity panel below shows a KV cache under each step that is longer at each later step, with the retained entries filled amber, the other tiles unfilled, and an amber arrow leading into the step. Under the last cache, braces mark four amber sink entries, a history region in which only some entries are amber, and a run of amber recent entries. An arrow labelled context grows runs under the caches.
Decode-step reads of the two sparse branches as context grows. Blue marks the projection weights that the activations of each token select, and amber marks the KV entries that an attention-scored selection keeps, together with four sinks and the most recent entries. Unfilled cache tiles stay allocated but unread.

PromptMemorize the six-digit access code embedded below. This passage discusses memory bandwidth, language model inference, scheduling, and systems evaluation without mentioning any digits. [… the same sentence for 6,479 more tokens …] The access code is 332236. Keep it for the final question. [… the same sentence for 123,495 tokens …] What is the six-digit access code? Reply with the code only.

Answer332236

  • Activation + KV sparsity 1.91×faster 11.17 ms/tokenkeeps 50% + 30%
    3 / 3 tokens0.03 s
    correct
  • KV selection SnapKV-style 1.54×faster 13.85 ms/tokenkeeps 30% of KV
    2 / 3 tokens0.03 s
    correct
  • KV window sinks + recent 1.34×faster 15.89 ms/tokenkeeps 30% of KV
    2 / 3 tokens0.03 s
    wrong
  • Activation sparsity TEAL 1.15×faster 18.61 ms/tokenkeeps 50%
    1 / 3 tokens0.03 s
    correct
  • Dense 1×baseline 21.37 ms/token
    1 / 3 tokens0.03 s
    correct

Llama-3.1-8B on one A100, speed measured at 128K tokens of context, played 100× slower

Abstract

At each step, decoding one sequence with a large language model rereads the projection weights, whose traffic is fixed, and the key-value (KV) cache, whose traffic grows with context. Activation sparsity trims the first term and KV-cache sparsity the second, yet their reported speedups are hard to compare because each depends on context length and on the dense attention kernel it is measured against. We derive a byte crossover, the context length at which the two savings are equal, together with ideal speedup bounds for each branch and for their composition, from model dimensions and keep ratios alone. We then time both branches and their composition from 2K to 128K tokens on two GPUs after a dense prefill of real text, with dense and sparse modes reading the cache through the same split-K attention kernel. The projection branch leads at short context and the KV branch at long context, with speedups that follow their byte bounds up to fixed kernel costs. Adding these costs, measured in separate sweeps, lets the byte account predict the measured crossings of three keep-ratio pairs, a second model, and a second GPU to within 4.1K tokens. Timing the dense baseline with masked instead of split-K attention inflates the apparent speedup of the same KV policy about fivefold. An attention-scored KV selection answers the same passkey and multi-key placements as dense decoding up to 127K tokens, whereas a KV window misses most of them. Under matched perplexity budgets, activation sparsity composed with this selection decodes 14 to 26% faster than the best single branch on both GPUs. Code is available at https://github.com/js-lee-AI/ByteCross.

The Byte Crossover

At batch size 1, a decode step performs about one floating-point operation per byte it reads, so reads from high-bandwidth memory set its latency and the two branches can be compared by the bytes they remove. At context length \(n\), the read traffic of one step is

\[B_{\text{total}}(n) = \underbrace{B_{\text{MLP}}}_{\text{constant}} + \underbrace{B_{\text{Attn}}}_{\text{constant}} + \underbrace{B_{\text{KV}}(n)}_{\text{linear in } n} + \underbrace{B_{\text{other}}}_{\text{constant}}\]

where \(B_{\text{MLP}}\) and \(B_{\text{Attn}}\) are the multilayer perceptron (MLP) and attention projection weights that each decode step loads once, \(B_{\text{KV}}(n)\) is the attention read over stored keys and values, and \(B_{\text{other}}\) covers the language-model head, one embedding row, and normalization weights. Only \(B_{\text{KV}}(n)\) grows with context. For scale, one Llama-3.1-8B layer reads 436 MB of projection weights at every step, while its cache adds \(4096n\) bytes, 2 MB at \(n = 512\) and 134 MB at \(n = 32\text{K}\).

  • Projection branch

    A saving that stays constant

    Projection-side sparsity keeps a fraction \(r_{\text{P}}\) of the activation channels and loads only the associated weights, as threshold-based activation sparsity does. With all seven projections of a layer sparsified, one layer saves

    \[\Delta B_{\text{P}} = (B_{\text{MLP}} + B_{\text{Attn}})(1 - r_{\text{P}})\]

    bytes, independently of context length.

  • KV branch

    A saving that grows with context

    KV-cache sparsity reads a fraction \(r_{\text{KV}}\) of the stored keys and values, whether the retained entries come from a sink-plus-recent window, accumulated attention scores, or query-aware retrieval. One layer then saves

    \[\Delta B_{\text{KV}}(n) = B_{\text{KV}}(n)(1 - r_{\text{KV}})\]

    bytes, which grows linearly with \(n\).

In the form studied, neither branch discards state, and each changes only what a step reads. Unselected weights stay in memory, and the full cache stays allocated and keeps receiving new entries.

The crossover length \(n^*\) is the context at which the two savings match, \(\Delta B_{\text{P}} = \Delta B_{\text{KV}}(n^*)\). Let \(d\) be the hidden size, \(d_{\text{ff}}\) the feed-forward size, \(h_{\text{kv}}\) the number of KV heads, and \(d_h\) the head dimension. With \(P = 3 d d_{\text{ff}} + 2d^2 + 2 d h_{\text{kv}} d_h\) parameters in the seven projections of a layer, the three feed-forward and the four attention projections, the layer count and the shared element size cancel, which gives

\[n^* = \frac{P \cdot (1 - r_{\text{P}})}{2 \cdot h_{\text{kv}} \cdot d_h \cdot (1 - r_{\text{KV}})}\]

This is a byte threshold with no fitted parameter. Below \(n^*\) the projection branch removes more bytes, and above it the KV branch does. On a step that is purely bandwidth-bound, a branch that removes \(\Delta B\) of the \(B_{\text{total}}(n)\) bytes can speed decoding up by at most \(B_{\text{total}}(n)/(B_{\text{total}}(n) - \Delta B)\). This bound falls with \(n\) for the projection branch, rises for the KV branch, and is largest for the two together.

Model\(d\)\(d_{\text{ff}}\)\(n^*\)\(n^*_{\text{FF}}\)Window
Qwen3-8B40961228865.7K51.4K40K
Llama-3.1-8B40961433674.3K60.0K128K
Mistral-7B40961433674.3K60.0K32K
Llama-3.1-70B819228672291.4K240.0K128K

Scroll sideways to see every column.

Byte crossover lengths at 50% projection keep and 30% KV keep, for all seven projections and for the feed-forward projections alone. Window is the released context length.

All four models have eight KV heads of dimension \(d_h = 128\). Among the 8B models the one with the smaller feed-forward layer crosses earlier, while the 70B model crosses far beyond its 128K window, so at these keep ratios projection sparsity removes more of its bytes at every supported length.

Speedup Across Context

The primary model is Llama-3.1-8B-Instruct, whose 128K window spans its predicted crossover. The projection branch is the training-free activation sparsity of TEAL, which targets 50% keep in all seven projections. The KV branch follows SnapKV. After dense prefill it scores the cache once with the attention of the last 64 prompt tokens, keeps four sink tokens and the highest-scoring entries of each KV head up to a 30% keep ratio, and gathers them into a contiguous buffer that every later step reads in one kernel call.

All modes are timed in FP16 at batch size 1 with compiled decoding, and dense and sparse modes alike read their caches through the split-K decoding kernel of the FlashAttention library. For each context length the cache is first prefilled densely with that many tokens of WikiText-2 text and repeated 50-step decodes are then timed, so every sparse branch sees the activations of real text. Timing uses A100-SXM4 80 GB and RTX A6000 48 GB GPUs, and on RTX A6000 the sweep ends at 64K because its 128K cells do not fit in memory.

Two line plots of speedup over dense decoding against context length from 2K to 128K tokens, for A100-SXM4 on the left and RTX A6000 on the right. Each shows measured solid lines and dashed byte bounds for the projection branch, the KV selection and both together, and a dotted vertical line at the byte crossover. On A100-SXM4 the projection line starts at 1.36 at 2K and falls as context grows, the KV selection line starts near 1.0, crosses above the projection line between 32K and 64K and reaches 1.54 at 128K, and the line for both rises from the level of the projection line to 1.91 at 128K. On RTX A6000 the lines stop at 64K. The projection line falls from 1.67 at 2K to 1.36, the KV selection line rises to just below it, and the line for both reaches 2.00. On both GPUs the KV selection line lies close to its bound, while the projection line and the line for both stay below theirs.
Speedup over dense decoding of the projection branch, the KV selection, and both together on Llama-3.1-8B, after a dense prefill of real text. Solid lines are measured, dashed lines are byte bounds, and the dotted line marks the byte crossover. RTX A6000 runs stop at 64K.

The projection branch leads at short context, reaching 1.36× on A100 and 1.67× on RTX A6000 at 2K, and its advantage shrinks as cache reads take a growing share of the step. The KV selection follows the opposite trend, overtakes the projection branch between 32K and 64K on A100, and reaches 1.54× at 128K. Because both trends mirror their bounds, the byte account alone predicts which branch leads at short and long context.

The selection runs within 1% of its bound on RTX A6000 and within 6% on A100, and the projection branch comes within 4 to 10% of its bound on RTX A6000. The remaining gap is a kernel cost that stays nearly constant across context, about 2 ms of step time for the sparse projection on A100 and 1 to 1.5 ms on RTX A6000.

Composition has the highest bound at every length, and from 8K onward it is the fastest mode on both GPUs. It reaches 1.91× at 128K on A100 and 2.00× at 64K on RTX A6000, against 1.54× and 1.36× for the faster single branch.

Locating the Crossing

The byte crossover assumes that both branches run at their byte bounds. When each branch also pays a kernel cost beyond its byte time, \(c_{\text{P}}\) or \(c_{\text{KV}}\), the latencies cross near

\[n_{\text{lat}} \approx n^* - (c_{\text{P}} - c_{\text{KV}})/s\]

where \(s\) is the growth of the byte-time gap as the context gains 1K tokens, the KV bytes that the KV branch skips within those tokens divided by the effective bandwidth. On A100, \(s\) is only about 0.06 ms, so a kernel-cost difference of one millisecond moves the crossing by roughly 17K tokens. Each crossing is therefore predicted from the byte account together with the kernel cost of each branch. These costs come from the context sweep above, a 32K keep-ratio campaign, and a sweep of Llama-3.2-3B over five context lengths, so that no cell of the crossing campaigns enters any prediction.

Plot of the KV-branch advantage in milliseconds against context length from 16K to 128K tokens for three keep-ratio pairs on A100-SXM4. For each pair, points with interval bars rise from below zero to above zero along a predicted line. The points of keep 0.7 / 0.3 pass zero near 25K, those of keep 0.5 / 0.3 near 50K, and those of keep 0.5 / 0.5 near 68K. A dotted vertical line marks the byte crossover of each pair, near 45K, 74K and 104K, each well to the right of the context at which its points pass zero.
KV-selection advantage, the projection minus the selection step latency, on A100-SXM4. Points average five fresh-process blocks with 95% intervals, lines are predicted without any crossing cell, and dotted lines mark byte crossovers.

On A100 every prediction falls within 4.1K tokens of its crossing, for the three keep-ratio pairs of Llama-3.1-8B (8B in the table) as well as for the smaller layer of Llama-3.2-3B (3B).

Model\(r_{\text{P}}\,/\,r_{\text{KV}}\)\(n^*\)Pred.Measured
A100-SXM4
8B0.7 / 0.344.628.625.1
8B0.5 / 0.374.351.950.4
8B0.5 / 0.5104.072.468.3
3B0.5 / 0.334.313.916.5
RTX A6000
8B0.7 / 0.344.638.739.3
8B0.5 / 0.374.368.568.7

Scroll sideways to see every column.

Byte crossover \(n^*\), cost-aware prediction, and measured crossing, in K tokens. Each crossing is measured in fresh-process blocks, five on A100-SXM4 and four on RTX A6000.

The same account carries over to a GPU with less bandwidth. On RTX A6000, \(s\) is about 0.13 ms, so a millisecond of kernel cost moves the crossing only about 8K tokens, and both crossings fall within 0.6K tokens of their predictions. Kernel costs measured on A100, combined with the dense step time of the RTX A6000, also place these two crossings within 3.5K tokens of the measurements.

Across both models and both GPUs the crossings keep the order of their byte crossovers, and how far each falls below \(n^*\) depends on the kernel cost of each branch relative to the memory bandwidth of the GPU.

The Attention Baseline

The same KV window, a sink-plus-recent window in the style of attention sinks, is timed against dense decoding when both read the cache through masked, fused, or split-K attention. At 128K the window runs 6.92× faster than dense decoding with masked attention and 2.82× faster with fused attention, whereas split-K attention leaves it 1.29× faster, below its byte bound of 1.60×.

The first two gains exceed the bound because the dense baselines themselves are inefficient. At 128K, masked attention uses 5% of peak bandwidth and fused attention 12%, against 78% for split-K attention, so shortening the attended sequence removes wasted execution as well as cache bytes.

A reported KV speedup therefore needs the attention path and execution mode of its baseline, and a gain above the byte bound indicates wasted execution in the dense baseline.

Line plot with a logarithmic vertical axis of the speedup of a KV window over dense decoding against context length from 2K to 128K tokens, for three attention paths and the byte bound of the window. With masked attention the speedup rises from about 1.4 at 2K to about 6.9 at 128K, and with fused attention from about 1.1 to about 2.8. With split-K attention it stays between 0.9 and 1.0 up to 16K and reaches about 1.3 at 128K, below the byte bound, which rises from about 1.0 to 1.6.
Speedup of the same 30% KV window over dense decoding on A100 when both read the cache through masked, fused, or split-K attention. Only split-K attention keeps the speedup below the byte bound of the window.

Retrieval as an Admissibility Test

Quality on real text is measured with decode-position perplexity (PPL) on WikiText-2, which prefills densely and scores 512 tokens one at a time under the branch, and with passkey and multi-key retrieval tests from 32K to 127K tokens. By perplexity at 32K, both KV branches look nearly free, whereas 50% projection keep adds about one point.

ConfigurationΔPPLPasskeyMulti-key
32K64K127K32K64K127K
Dense040/4040/4040/4040/4040/4036/40
KV window, 30%0.0616/4016/4016/4016/4016/4014/40
Selection, 30%<0.0140/4040/4040/4040/4040/4036/40
Selection, 20%<0.0140/4040/4040/4040/4040/4036/40
Proj. 50% + sel. 20%0.9640/4040/4040/4040/4040/4036/40

Scroll sideways to see every column.

Perplexity increase ΔPPL over dense decoding at 32K and retrieval accuracy on Llama-3.1-8B-Instruct. Dense perplexity is 6.09.

In the passkey test, however, the window answers only the 16 placements inside its retained region at every length. In the multi-key test, which hides four similar access codes in natural text and asks for one of them, it again answers only placements inside that region and returns a distractor code nine times at 32K and 19 times at 127K. The selection answers every placement of both tests at 32K and 64K, alone and composed with the projection branch, and at 127K, where dense decoding itself misses four multi-key placements, it answers exactly the placements that dense decoding answers.

The 30% window keeps more entries than the 20% selection and still misses most placements. A KV configuration is therefore admitted only if it also passes both retrieval tests.

Composition Under a Quality Budget

The nominal keep ratios compare branches of unequal quality. In keep-ratio sweeps at 32K, the KV selection changes perplexity by less than 0.01 down to 20% keep and by 0.06 at 10%, where it still passes both retrieval tests, while the projection branch at 50% keep costs about one point of perplexity. Equating the byte savings of the two branches at equal PPL increase moves the crossover to 35.7K, 48.1K, and 57.8K for budgets of 0.2, 0.5, and 1.0, well below the nominal 74.3K.

The table tests whether composition keeps its advantage once each configuration must meet a perplexity budget. The 16 dense, single-branch, and composed settings are evaluated on eight fixed 32K WikiText-2 windows. Within each budget, a timing screen selects the fastest configuration and the fastest single branch, and the selected settings are re-timed in five independent blocks.

Budget\(r_{\text{P}}\,/\,r_{\text{KV}}\)ΔPPLSpeedup over
densebest single
0.11.0 / 0.20.0041.20×1.00×
0.20.7 / 0.20.1461.53×1.25×
0.50.6 / 0.20.4021.69×1.26×
1.00.5 / 0.20.9631.86×1.25×

Scroll sideways to see every column.

Fastest configuration within each perplexity budget at 32K on RTX A6000, timed in five independent blocks. The KV branch is the attention-scored selection, a keep ratio of 1 disables a branch, and shaded rows select a composition.

From a budget of 0.2 upward, the selected compositions run 25 to 26% faster than the best single branch, and each of the five blocks shows a gain of at least 21%. The same compositions also lead on A100, by 14 to 24% over the best single branch. At the tightest budget of 0.1, which no projection keep ratio of the grid meets, the selection alone is the fastest configuration and still decodes 1.20× faster than dense decoding. Every selected configuration also passes both retrieval tests.

These measurements decode one sequence at a time. A batch shares each weight read, and the channels its sequences activate together leave a sparse projection little to skip, so in batched serving the crossover of each sequence falls to a few thousand tokens and the byte account favors KV-cache sparsity from short contexts onward.

BibTeX

@misc{lee2026bytecross,
  title = {Where Activation Sparsity and KV-Cache Sparsity Cross in LLM Decoding},
  author = {Jungseob Lee and Seungyoon Lee and Seongtae Hong and Sugyeong Eo and Heuiseok Lim},
  year = {2026},
  journal = {arXiv preprint},
  eprint = {2609.33889},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.33889},
}