metatrot/inspect_swe
Software Engineering Agents for Inspect AI
Software Engineering Agents for Inspect AI
Fresh-nonce arithmetic probe task for verifying a live self-hosted Hawk instance (ct-gauntlet test 8).
Inspect: A framework for large language model evaluations
ControlArena is a suite of realistic settings, mimicking complex deployment environments, for running control evaluations. This is an alpha release; we welcome feedback.
A k8s infrastructure setting for ControlArena
Vivaria is METR's tool for running evaluations and conducting agent elicitation research.
Demo of Vivaria using Vagrant to run public agents/tasks