signalrush/tau2-bench
τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Framework for evaluating and improving agents
Democratizing Reinforcement Learning for LLMs
LeWorldModel (PushT) - JEPA world model optimization hive task
Personal context vault: tree of markdown for knowledge, ideas, tasks. Claude Code skill + D3 visual server.
Python SDK that makes every Claude Code CLI tool callable as a Python function
Chronon is a data platform for serving for AI/ML applications.
Fork of instructkr/claude-code
A minimal primitive for self-controlling agents. step() + Python = every agent architecture pattern.
Train the smallest LM you can that fits in 16MB. Best model wins!