Repositories
james-aung-aisi/inspect_evals
Collection of evals for Inspect AI
james-aung-aisi/textquests
james-aung-aisi/mask
Code for evaluating AI systems on the MASK honesty benchmark.
james-aung-aisi/inspect_cyber
An Inspect extension for agentic cyber evaluations
james-aung-aisi/factorio-learning-environment
A non-saturating, open-ended environment for evaluating LLMs in Factorio
james-aung-aisi/ppo
Reimplementation of PPO with PyTorch
james-aung-aisi/qlearning
Reimplementation of Q-learning with a DQN
james-aung-aisi/risks
james-aung-aisi/reinforce
Reimplementation of REINFORCE
james-aung-aisi/ioccc-winner
Winners of the International Obfuscated C Code Contest
james-aung-aisi/control-arena
ControlArena is a collection of settings, model organisms and protocols - for running control experiments.
james-aung-aisi/inspect_ai
Inspect: A framework for large language model evaluations
james-aung-aisi/z2h
james-aung-aisi/evals
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
james-aung-aisi/SWELancer-Benchmark
This repo contains the dataset and code for the paper "SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?"
james-aung-aisi/mle-bench
MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering
james-aung-aisi/SWE-bench
james-aung-aisi/modular-public
james-aung-aisi/Remove-First-Score
Plugin for MuseScore, removes files
james-aung-aisi/bhp
james-aung-aisi/rl-algos
Reimplementations of some RL algos