Repositories
Felhof/DiscreteSAC
Felhof/sandbagging-eval-experiments
Felhof/ARENA_evals
Felhof/the-elicitation-game
Felhof/sandbagging
Felhof/DRL-from-Human-Preferences
A reproduction of the paper "Deep Reinforcement Learning from Human Preferences"
Felhof/steering-vectors
Steering vectors for transformer language models in Pytorch / Huggingface
Felhof/From-Sycophancy-To-Sandbagging
Felhof/TransformerLens
Felhof/PersonaInvestigations
Felhof/connectome
Felhof/Activation-Engineering-Investigations
Felhof/Deep-Reinforcement-Learning-Algorithms
Felhof/Comparing-Measures-of-LLM-Truthfulness
Felhof/LLM-Classification-Faithfulness
Felhof/Exhibiting-Deception-in-LLMs
Felhof/MLAB-Transformers-From-Scratch
Reimplementing transformers from scratch (from Redwood Research's Machine Learning for Alignment Bootcamp).
Felhof/DecisionTransformerInterpretability
Interpreting how transformers simulate agents performing RL tasks
Felhof/ARENA_2.0
Felhof/swap-graphs
An implementation of input swap graphs. A tool to discover the role of neural network components with causal interventions.
Felhof/pytest
The pytest framework makes it easy to write small tests, yet scales to support complex functional testing
Felhof/MEng_Project
Felhof/Kaggle_Houseprices
Felhof/kaggle_titanic
Felhof/Pintos
Implementation of scheduler and user programs for the Pintos Operating System - Imperial College 2nd Year Lab
Felhof/Wacc-Compiler
Compiler for the WACC language specified in Imperial College 2nd Year Compilers course