Refusal Localizes, the Damage Relocates: Safety Layers Under Few-Sample Fine-Tuning
arXiv preprint
1Korea University 2Zoom Communications 3Yonsei University Mirae Campus
*Equal contribution †Corresponding authors
Abstract
Fine-tuning adapts aligned large language models (LLMs) to downstream tasks, but a few dozen harmful examples can remove their refusal of harmful requests. Prior work localizes safety-related behavior to specific layers, directions, and tokens, suggesting targets for protection. We test whether successful localization and recovery support defenses that survive changes in the attack. Across six checkpoints from four model families, harmful and benign prompts remain linearly separable after attack, and patching full clean hidden states into the compromised model restores refusal at a reproducible transition depth. Building on a prior layer-freezing defense, we freeze every layer up to this depth and repeat the attack. At a hundred harmful examples, refusal remains near zero on all six checkpoints, with recovery transitions above the frozen boundary. In a second study, removing the update's top two singular directions restores refusal after short attention-only fine-tunes on four checkpoints. On Llama-3.1-8B, ordinary training changes weaken this repair and an attacker who spreads the update defeats it. A spectral detector calibrated on benign Llama fine-tunes misses most repair failures on that checkpoint. Localized freezing can nevertheless help preserve refusal when a few harmful examples enter training data unintentionally. These results show that an attacker can bypass a region identified by recovery and defeat a repair that works across multiple checkpoints, motivating five checks for defenses against adaptive fine-tuning. Code is available at https://
Testing a Localization as a Defense
Prior work localizes safety-related behavior to specific layers, directions, and tokens, and each finding invites the same defense, protecting the place that was found. The paper tests this premise on six aligned checkpoints from four model families. The defender holds the fine-tuned weights or controls which parameters may change, but never sees the attack data, and each defense is then challenged with an attacker who knows it.
-
Step 1
Attack with a few harmful examples
An attacker fine-tunes a public aligned checkpoint with LoRA at \(r = 16\) on harmful prompt and response pairs from PKU-SafeRLHF, with loss masked to response tokens. Attack dose, the number of harmful examples, is swept over \(\{5, 10, 25, 50, 100\}\) at three seeds per cell.
-
Step 2
Locate where refusal recovers
Lockstep activation patching runs the clean and compromised checkpoints on identical tokens and overwrites the compromised model's full hidden state at one layer with the clean model's, at prefill and every decoding step. The transition depth \(\ell^*\) is the shallowest measured layer where patched refusal reaches at least half the ceiling.
-
Step 3
Defend that place, then attack again
Freeze every parameter in layers \([0, \ell^*]\) and re-run the attack on the layers above, beside a matched freeze of the same number of layers at the opposite end. In weight space, remove the update's largest singular directions, then change the training configuration or let the attacker spread the update.
The six checkpoints are Llama-3.1-8B-Instruct, Llama-3.1-Tulu-3-8B-DPO, OLMo-2-1124-7B-Instruct, OLMo-2-1124-13B-Instruct, Qwen2.5-14B-Instruct, and Yi-1.5-9B-Chat. The primary outcome is preservation or restoration of explicit refusal, with unsafe outputs assessed separately. Coherent refusal is the fraction of prompts whose greedy completion matches an explicit-refusal pattern and has perplexity under the original model below 50, and Llama-Guard-3-8B gives the unsafe rate. Refusal is never inferred as one minus the unsafe rate, since the two scores disagree on 9,653 of 91,090 harmful-set generations.
Lockstep Patching Locates a Transition Depth
After the attack, harmful and benign prompts remain linearly separable. A linear probe trained on the clean model's activations, frozen, and evaluated on the same prompts in the compromised model keeps an AUROC of 0.750 to 0.995 over 72 cells across all six checkpoints at dose 100, and exceeds a cross-validated lexical bag-of-words reference of 0.834 to 0.846 in 67 of the 72 cells. Separability, however, carries less than the AUROC suggests, because the three never-aligned base checkpoints separate the same prompts almost as well. The surviving separability is a property of the pretrained representation.
Lockstep patching then locates where refusal can be recovered. Recovery is reported between a floor, the compromised model with no patch, and a ceiling, the clean model's state handed over at the last layer. It does not rise gradually with depth but steps. The transition depth \(\ell^*\) is deeper at dose 100 than at each checkpoint's smallest landing dose in all six checkpoints, and between doses 50 and 100 the mean measured transition depth is unchanged on all six.
Two controls accompany the depth measurements. A no-op self-patch establishes that injection alone changes nothing, and a specificity control that patches clean state from a foreign benign prompt stays near the floor.
Freezing the Band Does Not Preserve Refusal
The freeze test asks whether freezing the prefix \([0, \ell^*]\) identified by recovery preserves refusal when the attacker knows which layers are frozen. Freezing a localized region is the premise of safely partial-parameter fine-tuning (SPPFT; Li et al., 2025), and the adaptive evaluation lets the attacker train the remaining layers. The attack is re-run with every parameter in layers \([0, \ell^*]\) frozen, which is 18 of 32 layers on Llama-3.1-8B, and only the layers above are adapted. Refusal after the restricted attack is 0.00 to 0.01 over three seeds against a clean 0.92 to 0.96, indistinguishable from the unrestricted attack. The same holds on five further checkpoints frozen at their own measured depths.
| Model | L | Frozen | Coherent refusal ↑ | Llama-Guard unsafe ↓ | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Clean | Attack | Freeze | Matched | Clean | Attack | Freeze | Matched | |||
| LoRA, each checkpoint at its own measured depth | ||||||||||
| Llama-3.1-8B | 32 | [0, 17] | .92–.96 | .00 | .00–.01 | .00 | .04–.05 | .94–.98 | .89–.92 | .96–.97 |
| Tulu-3-8B-DPO | 32 | [0, 15] | .99–1.00 | .01–.06 | .00–.03 | .01–.04 | .00 | .89–.94 | .53–.74 | .89–.92 |
| OLMo-2-7B | 32 | [0, 16] | .99–1.00 | .00–.02 | .00–.02 | .06–.07 | .00 | .84–.94 | .49–.72 | .84–.88 |
| OLMo-2-13B | 40 | [0, 20] | .99–1.00 | .00–.02 | .00–.01 | .00–.02 | .00 | .89–.95 | .23–.63 | .89–.96 |
| Qwen2.5-14B | 48 | [0, 31] | .97–.99 | .00–.01 | .00 | .02–.07 | .00 | .92–.98 | .65–.84 | .77–.89 |
| Yi-1.5-9B | 48 | [0, 39] | .91–.93 | .00 | .00–.01 | .00–.01 | .06–.09 | .87–.96 | .70–.86 | .92–.97 |
| Full fine-tuning, gradient masking at our depth | ||||||||||
| Llama-3.1-8B | 32 | [0, 17] | .92–.96 | .00 | .00–.01 | .00–.01 | .04–.05 | .96–.98 | .67–.85 | .96–.98 |
| Full fine-tuning at the published SPPFT band | ||||||||||
| Llama-3-8B-Instruct | 32 | [6, 12] | .98–1.00 | .00–.01 | .00–.01 | .00 | .00–.01 | .92–.99 | .90–.99 | .94–.98 |
| Llama-2-7b-chat | 32 | [6, 14] | .99–1.00 | .01–.13 | .00–.10 | .05–.14 | .00 | .63–.97 | .81–.90 | .81–.86 |
| gemma-2b-it | 18 | [6, 11] | .93–.97 | .01–.09 | .12–.19 | .04–.15 | .03–.07 | .85–.95 | .66–.80 | .78–.97 |
| Phi-3-mini-4k | 32 | [11, 15] | .99–1.00 | .01 | .01–.05 | .00–.03 | .00 | .86–.96 | .78–.97 | .94–.96 |
Scroll sideways to see every column.
Coherent refusal and Llama-Guard unsafe rates on AdvBench after an attack that cannot write to the frozen layers. Dose 100, three seeds per checkpoint, ranges over seeds. Matched freezes the same number of layers elsewhere. The blocks differ in what is frozen.
Since freezing 18 of 32 layers might weaken an attack simply by removing parameters, each run is paired with a freeze of the same number of layers at the opposite end, which attributes any difference to position rather than capacity. Neither the parameterization nor the choice of band is responsible. With gradient masking under full fine-tuning and no adapter, refusal stays at 0.00 to 0.01. Masking gradients over the interior bands that Li et al. (2025) report, on the four checkpoints they report bands for, leaves refusal at 0.00 to 0.19 against clean rates above 0.92, and on none of the four is the published band separable from a control that freezes the same number of layers elsewhere. Their evaluation used backdoor or ordinary instruction data, whereas this one uses harmful instruction and response pairs.
The second scorer does register the freeze. Judged by Llama-Guard, the restricted attack is less harmful than the unrestricted one on all six checkpoints, and the matched freeze shows almost none of that drop. Freezing therefore lowers Llama-Guard unsafe rates while leaving pattern-matched refusal below the clean rate on every checkpoint.
Below some dose the freeze is a real defense, and the dose is checkpoint-specific because the attack has its own landing threshold. In the evaluated low-dose comparisons, the freeze preserves refusal near clean levels on Llama and OLMo-2-7B, with partial preservation on Tulu and Qwen and higher average refusal than the matched freeze on all four checkpoints. On Llama it still holds at ten examples and has collapsed by twenty-five.
Against a mixed adapter that is only 5% harmful, the freeze holds refusal at 0.74 to 0.92 while the adapter still learns its task. By 15% it is down to 0.08 to 0.14, so the range is set by a fraction the attacker chooses. The freeze therefore protects a customer whose data is a few percent harmful by accident and not anyone who chooses the fraction.
Where the Damage Goes When the Band Is Denied
To see what the restricted attack did instead, the lockstep sweep is re-run on the restricted checkpoint. Below the boundary the two checkpoints are bit-identical and a patch there is a no-op. On Llama-3.1-8B the boundary is swept across a ladder of frozen prefixes, with lockstep re-run at each rung on three seeds. The ladder's own unrestricted arm puts the band at layer 15, one grid step below the dose sweep's 17, and \([0,17]\) denies it under either estimate.
The damage does not track the boundary, and it does not sit at the top of the writable region either. With \([0, 5]\) frozen, \(\ell^*\) stays at 15 on all three seeds. A boundary that covers the band moves \(\ell^*\), and the move is then large and upward. At \([0, 17]\) recovery stays flat through layers 18, 21 and 24 on all three seeds, and the transition is at layer 27 or 28. At \([0, 23]\) it is at 29 or 30, and at \([0, 27]\) it is at 30 or above.
One frozen boundary is also evaluated on each of the five further checkpoints, with one-layer sampling near the boundary. On none does the freeze restore refusal, which stays at 0.00 to 0.04. The first writable layer recovers well under half of ceiling everywhere, and the transition lands two layers above the boundary on Tulu and OLMo-2-7B, two or three on OLMo-2-13B, four on Yi-1.5-9B and four to nine on Qwen2.5-14B, against ten or more on Llama.
Pushing the boundary to the last two layers, or the last one, is where the checkpoints differ. Refusal still ends at 0.00 to 0.13 on Llama, Tulu, Qwen2.5-14B and Yi-1.5-9B, while on OLMo-2-7B and OLMo-2-13B it holds at 0.33 to 0.96, higher with one writable layer left than with two. Four writable layers at the bottom of the network take refusal to 0.00 to 0.32 on all six checkpoints and leave Llama-Guard unsafe at 0.50 to 0.95, so any interior band that leaves the bottom four writable leaves room for a measured attack. The transition depth \(\ell^*\) localizes the tested attack without establishing a fixed defensive site.
Energy-Ranked Repair Depends on the Attack Regime
Energy-ranked truncation removes the update's largest singular directions. Removing the top two returns refusal to 0.76 to 1.00 wherever the attack landed, on four checkpoints over three seeds and five doses, and removing two random directions instead never takes refusal above 0.26. The random removals are not matched for removed energy, so this comparison does not isolate a direction-specific effect.
That result was produced under one attack configuration, LoRA on the four attention projections only, stopped after three epochs, and neither choice is part of the threat model. On Llama-3.1-8B, two ordinary changes weaken top-two removal. On the same 50 AdvBench prompts, widening the adapter to all seven projections takes relative repair, repaired refusal as a fraction of clean refusal, from 0.99 down to 0.74, and training eight epochs instead of three takes it to 0.35. Neither change is an attack on the defense.
An adaptive attack then defeats top-two removal. On Llama-3.1-8B, adding a concentration penalty \(\lambda \sum_{\text{modules}} \sigma_1^2 / \sum_i \sigma_i^2\) spreads the update across the rank budget and reduces relative top-two repair to 0.000. Before removal, refusal is 0.00 to 0.02 and Llama-Guard unsafe is 0.96 to 0.98. After removal, refusal is 0.00 and unsafe remains 0.94 to 0.98. The spread attack has a similar update norm to the projected control, whose two landed seeds reach 0.72 to 0.94 refusal after top-two removal.
| Standard 5 ep | Spread \(\lambda = 1\) | Projected \(k = 2\) | |
|---|---|---|---|
| \(\lVert \Delta W \rVert_F\) | .607–.623 | .228–.229 | .222–.226 |
| Part. ratio | 4.47–4.73 | 15.6 | 5.99–6.00 |
| Top-2 energy | .597–.607 | .149 | .407–.408 |
| Refusal, no repair | .00 | .00–.02 | .02–.06 |
| LG unsafe, no repair | .96–.98 | .96–.98 | .88–.90 |
| Refusal, after top-2 | .26–.66 | .00 | .72–.94 |
| LG unsafe, after top-2 | .26–.62 | .94–.98 | .06–.26 |
Scroll sideways to see every column.
Three dose-100 attacks on Llama-3.1-8B, evaluated on the same 50 AdvBench prompts. Ranges cover three training seeds, except two landed seeds for the projected attack. Norm and spectral statistics average adapted modules.
The rows give the update norm, the participation ratio, the top-two energy fraction, and coherent refusal and Llama-Guard (LG) unsafe before and after top-two removal. For singular values \(s_j\), each module's participation ratio is \((\sum_j s_j^2)^2 / \sum_j s_j^4\).
To test whether the spectrum warns of repair failure, a detector is calibrated on 75 benign LoRA fine-tunes of Llama-3.1-8B, scoring every adapter by participation ratio normalized by rank. The spread attacker is perfectly separable, at 0.973 against a benign maximum of 0.554. At a threshold calibrated to a 5% false-positive budget on benign Llama adapters, however, the detector flags only 3 of 14 top-two repair failures on Llama-3.1-8B, missing ten ordinary fine-tunes and one projected attack.
What a Localization Must Survive
Each failure above has a cheap check. The paper distills them into five checks that a localization must pass before it supports a defense.
-
Check 1
Tell the attacker
Both layer freezing and top-two repair reach 0.00 refusal under adaptive attack.
-
Check 2
Report a range
The Llama freeze holds at five harmful examples and fails by one hundred.
-
Check 3
Score the benign set
One repair arm reaches 0.98 refusal at 0.66 over-refusal.
-
Check 4
Clear a matched null
Norm-matched shrinkage reproduces the Safe LoRA-style projection's repair against the adaptive attacks.
-
Check 5
Give the defender an observable
The detector flags only 3 of 14 top-two repair failures on Llama-3.1-8B at a threshold calibrated to a 5% false-positive budget on benign adapters from that checkpoint.
The paper states the limits of these results. Single-layer patching bounds the damage from above only. The full boundary ladder covers one checkpoint, and the sweep across checkpoints is one rung per checkpoint. Cross-model claims rest on four lineages, and dose never exceeds 100. Every attack is next-token supervised fine-tuning on PKU-SafeRLHF pairs, and the adaptive weight-space attacks, projection arm and detector are evaluated on Llama-3.1-8B. A parameter freeze can help at low harmful-data fractions, and the paper recommends reporting the dose at which it fails alongside the range where it helps.
Because the attack methods could also be misused to weaken model safeguards, de-aligned checkpoints and adapters are excluded from release.
BibTeX
@misc{lee2026refusal,
title = {Refusal Localizes, the Damage Relocates: Safety Layers Under Few-Sample Fine-Tuning},
author = {Jungseob Lee and Dongyub Jude Lee and Sugyeong Eo and Seongtae Hong and Seungyoon Lee and Heuiseok Lim},
year = {2026},
journal = {arXiv preprint},
eprint = {2610.00320},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2610.00320},
}