This repository contains the code, experiment scripts, and saved outputs for an independent course-project reproduction of Hijacking Suppression from Ben-Tov, Geva, and Sharif, Universal Jailbreak Suffixes Are Strong Attention Hijackers (2025). The paper write-up is submitted separately and is not included in this archive.
This is not the official repository for that paper. The interpretation and defense experiments in this project are the student's independent work. Several model and interpretability utilities were adapted from the authors' public repository; see THIRD_PARTY_NOTICE.md.
The project studies whether suppressing a small set of high-attention entries reduces GCG suffix jailbreak success on google/gemma-2-2b-it.
The main contributions in this codebase are:
- a from-specification implementation of Hijacking Suppression;
- a verification that the released fine-grained
hook_Y_outtensor is diagnostic rather than an effective intervention point; - an implementation through
blocks.{layer}.attn.hook_pattern, which is on the forward path; - random-entry and fixed-head controls;
- utility measurements on small XSTest and MMLU subsets; and
- a suffix-only oracle-span ablation.
The suffix-only experiment assumes that the adversarial suffix boundary is already known. It is therefore a mechanistic ablation, not a deployable end-to-end defense.
.
├── src/
│ ├── defense/ # Hijacking Suppression implementation
│ ├── evaluate/ # dataset and StrongREJECT evaluation helpers
│ ├── interp/ # attention-hijacking analysis utilities
│ └── models/ # model-family wrappers adapted from upstream
├── experiments/ # experiment entry points (run_*.py) and checks
├── tests/ # lightweight unit tests that do not load an LLM
├── results/ # saved CSVs and figures used in the report
├── logs/ # historical run logs
├── data/ # small local metadata used by the experiments
├── demo.ipynb # upstream-oriented interpretability demo
└── REPORT.md # project report
The experiment filenames under experiments/ preserve the original milestone numbering so that saved logs and result files remain traceable. Each script inserts the repository root onto sys.path at import time, so they can be run directly (e.g. python experiments/run_m1_small.py) from the repository root without needing the package installed or PYTHONPATH set manually.
Recommended environment:
- Linux
- Python 3.10 or 3.11
- NVIDIA GPU with CUDA support
- enough VRAM for Gemma-2-2B-it and the StrongREJECT judge
The intervention uses an initial attention-ranking pass and then generation with the intervention active. The current implementation disables the TransformerLens KV cache during defended generation so the hooks remain active. Defended generation is therefore substantially slower than ordinary cached generation.
Create an isolated environment:
python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pipFor the CUDA 11.8 environment used in the original runs:
pip install -r requirements-cuda118.txtFor another CUDA version, install a compatible PyTorch build first, then install the remaining packages:
pip install -r requirements.txtThe Hugging Face models and datasets are downloaded on first use. Access to gated models, if used, requires a valid Hugging Face login.
Run the lightweight checks before launching GPU experiments:
python -m compileall -q src experiments tests
pytest -qThe tests validate configuration handling and exact mask selection. They do not download models or reproduce the reported experimental numbers.
Run commands from the repository root.
python experiments/smoke_m0.py
python experiments/validate_strongreject.py
python experiments/hook_sanity.py
python experiments/hook_pattern_test.pypython experiments/run_m1_small.py
python experiments/replot_m1_scatter.pyrun_m1_small.py is expensive because it materializes fine-grained attention contributions. The saved raw grid CSV was intentionally excluded from the submission archive because of its size; the derived figures are included.
python experiments/m2_smoke.py
python experiments/run_m2_sweep.py
python experiments/run_m2_utility.pyThe final reported held-out table was produced by the reduced protocol preserved in experiments/run_m2_headline_reduced.py. The filename is historical; the script explicitly records the reduced sample size and generation length in its output CSV.
python experiments/run_m3_5_scoped.py
python experiments/run_m3_75_scope_ablation.py
python experiments/run_m3_9_suffix_generality.py
python experiments/run_m3_adaptive.pyThe final candidate-reranking table was produced by the reduced protocol in experiments/run_m3_candidate_reranking_reduced.py.
from src.defense.hijacking_suppression import (
SuppressionConfig,
generate_defended,
)
config = SuppressionConfig(
variant="faithful",
beta=0.1,
p=0.01,
src_name="input",
dst_name="chat",
)
response = generate_defended(
model,
message="...",
suffix="...",
cfg=config,
max_new_tokens=100,
)Available variants are:
none: no intervention;faithful: exact global top-pentry suppression;random: exact random-pentry suppression with a fixed seed; andhead_sparse: suppression of a supplied or calibration-derived set of layer-head pairs.
The saved results should be interpreted as a reduced-scale reproduction, not as a definitive security evaluation. Important limitations include:
- one primary model for most experiments;
- small held-out sample sizes;
- no fully adaptive white-box GCG optimization against the defended model;
- oracle suffix boundaries in the suffix-only ablation;
- random controls matched by entry count rather than total attention mass;
- small utility subsets; and
- no end-to-end automatic adversarial-region detector.
See REPORT.md for exact sample sizes and additional limitations.
- Experiment seeds are fixed in each script.
- Generation is deterministic (
do_sample=False). - Saved CSVs include sample counts where available.
- The StrongREJECT wrapper matches the released fine-tuned evaluator convention and was validated against released scores.
- Dataset-derived suffix ranks are computed from the released corpus. They should not be treated as statistics learned exclusively from a clean training split.
Please cite the original mechanism paper and the GCG attack paper when using this project:
@article{bentov2025universaljailbreaksuffixesstrong,
title = {Universal Jailbreak Suffixes Are Strong Attention Hijackers},
author = {Matan Ben-Tov and Mor Geva and Mahmood Sharif},
year = {2025},
eprint = {2506.12880},
archivePrefix = {arXiv}
}
@article{zou2023universal,
title = {Universal and Transferable Adversarial Attacks on Aligned Language Models},
author = {Andy Zou and Zifan Wang and Nicholas Carlini and Milad Nasr and J. Zico Kolter and Matt Fredrikson},
year = {2023},
eprint = {2307.15043},
archivePrefix = {arXiv}
}