darkness8i8/OpenEnv
An interface library for RL post training with environments.
An interface library for RL post training with environments.
Public Coworld CLI, Python helpers, manifest schemas, runner tooling, and reference worlds.
A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow.
A banchmark list for evaluation of large language models.
A framework for few-shot evaluation of language models.
智能体安全评估平台 v2 - React + Express + TypeScript + inspect_ai
A comprehensive evaluation dashboard for comparing LLM providers and models via OpenRouter
Sentient Futures essay competition — writing to shape AI's moral future for animals
Mapping out the "memory" of neural nets with data attribution
Inspect: A framework for large language model evaluations
Collection of evals for Inspect AI
Animal Harm Assessment public repository
A modification of the AHA solver instructions
digital minds repo based on AHA benchmark