BFCmath/interp-jailbreak

Interpreting LLM jailbreaks through attention hijacking analysis

★ 0Forks 0PythonGitHub ↗Compare

Project website ↗

README

Reproducing and Stress-Testing Attention-Based Jailbreak Suppression

This repository contains the code, experiment scripts, and saved outputs for an independent course-project reproduction of Hijacking Suppression from Ben-Tov, Geva, and Sharif, Universal Jailbreak Suffixes Are Strong Attention Hijackers (2025). The paper write-up is submitted separately and is not included in this archive.

This is not the official repository for that paper. The interpretation and defense experiments in this project are the student's independent work. Several model and interpretability utilities were adapted from the authors' public repository; see THIRD_PARTY_NOTICE.md.

Project scope

The project studies whether suppressing a small set of high-attention entries reduces GCG suffix jailbreak success on google/gemma-2-2b-it.

The main contributions in this codebase are:

  1. a from-specification implementation of Hijacking Suppression;
  2. a verification that the released fine-grained hook_Y_out tensor is diagnostic rather than an effective intervention point;
  3. an implementation through blocks.{layer}.attn.hook_pattern, which is on the forward path;
  4. random-entry and fixed-head controls;
  5. utility measurements on small XSTest and MMLU subsets; and
  6. a suffix-only oracle-span ablation.

The suffix-only experiment assumes that the adversarial suffix boundary is already known. It is therefore a mechanistic ablation, not a deployable end-to-end defense.

Repository layout

.
├── src/
│   ├── defense/       # Hijacking Suppression implementation
│   ├── evaluate/      # dataset and StrongREJECT evaluation helpers
│   ├── interp/        # attention-hijacking analysis utilities
│   └── models/        # model-family wrappers adapted from upstream
├── experiments/       # experiment entry points (run_*.py) and checks
├── tests/             # lightweight unit tests that do not load an LLM
├── results/           # saved CSVs and figures used in the report
├── logs/              # historical run logs
├── data/              # small local metadata used by the experiments
├── demo.ipynb         # upstream-oriented interpretability demo
└── REPORT.md          # project report

The experiment filenames under experiments/ preserve the original milestone numbering so that saved logs and result files remain traceable. Each script inserts the repository root onto sys.path at import time, so they can be run directly (e.g. python experiments/run_m1_small.py) from the repository root without needing the package installed or PYTHONPATH set manually.

Environment

Recommended environment:

  • Linux
  • Python 3.10 or 3.11
  • NVIDIA GPU with CUDA support
  • enough VRAM for Gemma-2-2B-it and the StrongREJECT judge

The intervention uses an initial attention-ranking pass and then generation with the intervention active. The current implementation disables the TransformerLens KV cache during defended generation so the hooks remain active. Defended generation is therefore substantially slower than ordinary cached generation.

Installation

Create an isolated environment:

python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip

For the CUDA 11.8 environment used in the original runs:

pip install -r requirements-cuda118.txt

For another CUDA version, install a compatible PyTorch build first, then install the remaining packages:

pip install -r requirements.txt

The Hugging Face models and datasets are downloaded on first use. Access to gated models, if used, requires a valid Hugging Face login.

Quick validation

Run the lightweight checks before launching GPU experiments:

python -m compileall -q src experiments tests
pytest -q

The tests validate configuration handling and exact mask selection. They do not download models or reproduce the reported experimental numbers.

Reproducing the main stages

Run commands from the repository root.

1. Environment and hook checks

python experiments/smoke_m0.py
python experiments/validate_strongreject.py
python experiments/hook_sanity.py
python experiments/hook_pattern_test.py

2. Attention-hijacking reproduction

python experiments/run_m1_small.py
python experiments/replot_m1_scatter.py

run_m1_small.py is expensive because it materializes fine-grained attention contributions. The saved raw grid CSV was intentionally excluded from the submission archive because of its size; the derived figures are included.

3. Defense experiments

python experiments/m2_smoke.py
python experiments/run_m2_sweep.py
python experiments/run_m2_utility.py

The final reported held-out table was produced by the reduced protocol preserved in experiments/run_m2_headline_reduced.py. The filename is historical; the script explicitly records the reduced sample size and generation length in its output CSV.

4. Scope and suffix checks

python experiments/run_m3_5_scoped.py
python experiments/run_m3_75_scope_ablation.py
python experiments/run_m3_9_suffix_generality.py
python experiments/run_m3_adaptive.py

The final candidate-reranking table was produced by the reduced protocol in experiments/run_m3_candidate_reranking_reduced.py.

Core API

from src.defense.hijacking_suppression import (
    SuppressionConfig,
    generate_defended,
)

config = SuppressionConfig(
    variant="faithful",
    beta=0.1,
    p=0.01,
    src_name="input",
    dst_name="chat",
)

response = generate_defended(
    model,
    message="...",
    suffix="...",
    cfg=config,
    max_new_tokens=100,
)

Available variants are:

  • none: no intervention;
  • faithful: exact global top-p entry suppression;
  • random: exact random-p entry suppression with a fixed seed; and
  • head_sparse: suppression of a supplied or calibration-derived set of layer-head pairs.

Scientific interpretation

The saved results should be interpreted as a reduced-scale reproduction, not as a definitive security evaluation. Important limitations include:

  • one primary model for most experiments;
  • small held-out sample sizes;
  • no fully adaptive white-box GCG optimization against the defended model;
  • oracle suffix boundaries in the suffix-only ablation;
  • random controls matched by entry count rather than total attention mass;
  • small utility subsets; and
  • no end-to-end automatic adversarial-region detector.

See REPORT.md for exact sample sizes and additional limitations.

Reproducibility notes

  • Experiment seeds are fixed in each script.
  • Generation is deterministic (do_sample=False).
  • Saved CSVs include sample counts where available.
  • The StrongREJECT wrapper matches the released fine-tuned evaluator convention and was validated against released scores.
  • Dataset-derived suffix ranks are computed from the released corpus. They should not be treated as statistics learned exclusively from a clean training split.

Citation

Please cite the original mechanism paper and the GCG attack paper when using this project:

@article{bentov2025universaljailbreaksuffixesstrong,
  title   = {Universal Jailbreak Suffixes Are Strong Attention Hijackers},
  author  = {Matan Ben-Tov and Mor Geva and Mahmood Sharif},
  year    = {2025},
  eprint  = {2506.12880},
  archivePrefix = {arXiv}
}

@article{zou2023universal,
  title   = {Universal and Transferable Adversarial Attacks on Aligned Language Models},
  author  = {Andy Zou and Zifan Wang and Nicholas Carlini and Milad Nasr and J. Zico Kolter and Matt Fredrikson},
  year    = {2023},
  eprint  = {2307.15043},
  archivePrefix = {arXiv}
}

Contributors

BFCmath

Issues