Repositories
DalasNoin/wellbeing
Measuring and improving the functional pleasure and pain of AIs
DalasNoin/open_whisper
DalasNoin/safety-tooling
Inference API for many LLMs and other useful tools for empirical research
DalasNoin/aiphishing_org
Website project for AI phishing
DalasNoin/cot_monitoring_environment
This environment should serve as a simple place to monitor the CoT of a model
DalasNoin/DeepResearch
Tongyi Deep Research, the Leading Open-source Deep Research Agent
DalasNoin/non-adversarial-reproduction
Official code for "Measuring Non-Adversarial Reproduction of Training Data in Large Language Models" (https://arxiv.org/abs/2411.10242)
DalasNoin/redteaming
redteaming a simple language model like gpt2. based on anthropic redteaming paper
DalasNoin/inspect_ai
Inspect: A framework for large language model evaluations
DalasNoin/control-arena
ControlArena is a suite of realistic settings, mimicking complex deployment environments, for running control evaluations. This is an alpha release; we welcome feedback.
DalasNoin/langchain-chainlit-docker-deployment-template
A template to run Lanchain Powered App using Chainlit Front UI
DalasNoin/anthropic-quickstarts
A collection of projects designed to help developers quickly get started with building deployable applications using the Anthropic API
DalasNoin/safety_benchmarks
Safety Benchmarks such as Refusal Bench
DalasNoin/refusal_direction
Code and results accompanying the paper "Refusal in Language Models Is Mediated by a Single Direction".
DalasNoin/LM-exp
LLM experiments done during SERI MATS - focusing on activation steering / interpreting activation spaces
DalasNoin/ActivationDirectionAnalysis
DalasNoin/aneurysm-segment
DalasNoin/weblm
Drive a browser with a language model
DalasNoin/exploring_modelgraded_evaluation
exploring model-graded evaluation
DalasNoin/GPTQ-for-LLaMa
4 bits quantization of LLaMA using GPTQ
DalasNoin/TextWorld
TextWorld is a sandbox learning environment for the training and evaluation of reinforcement learning (RL) agents on text-based games.
DalasNoin/chat-langchain
DalasNoin/langchain
⚡ Building applications with LLMs through composability ⚡
DalasNoin/simple-llama-finetuner
Simple UI for LLaMA Model Finetuning
DalasNoin/DecisionTransformerInterpretability
Interpreting how transformers simulate agents performing RL tasks
DalasNoin/Minigrid
Simple and easily configurable grid world environments for reinforcement learning
DalasNoin/reference_chatbot
In-Context Retrieval-Augmented Language Models AI21labs Implementation
DalasNoin/SVDInterpretTransformer
Apply SVD to Transformer weights