Kimiyu-186/shellbench
The agent benchmark that scores the full stack — harness, config, and model — not just the LLM. Trace-based scoring, reliability metrics, configuration diagnostics.
Gamer
The agent benchmark that scores the full stack — harness, config, and model — not just the LLM. Trace-based scoring, reliability metrics, configuration diagnostics.
Lightweight coding agent that runs in your terminal
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞