OliverLeeXZ/SERA
Official implement on "Self-Evaluating Recursive Agents"
Setbacks and farewells are but embellishments in life.
Official implement on "Self-Evaluating Recursive Agents"
Official implement on 'What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents'
[ACL 2026] OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces
[ICML 2026] Official implement on 'Beyond Mode Collapse: Distribution Matching for Diverse Reasoning'
[ACL 2026] Official implement on 'Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs'
A set of examples based on verl for end-to-end RL training recipes.
Open-source evaluation toolkit of large vision-language models (LVLMs), support 160+ VLMs, 50+ benchmarks
OpenClaw-RL: Train any agent simply by talking
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework