ArthurConmy/sae

★ 7Forks 0PythonGitHub ↗Compare

README

A quick SAE implementation

Maybe TODO for improvements:

  • Geometric median crap?
  • Untie the bias in and bias out?
  • Should we really be doing bias initialized to 0 in reinits? Seems like it makes a lot of non-zero fires, eek (Note: this didn't seem promising at work)
  • Should we Adam reinitialize better? ccLeo
  • Can't we make the neuron resampling procedure stochastic? i.e resample a neuron with some probability that's a function of how few times it fired
  • Should we improve the lib and make pip install git+ actually work from colab? Idk why it don't work ... then we sure should implement some good vizualization utils

Help

pip install -e . is optimal for now. torch >= 2.6 is required: older releases can run arbitrary code from torch.load (CVE-2025-32434), which this repo relies on to load checkpoints from wandb and the Hugging Face hub safely.

Contributors

ArthurConmyclaude

Issues