MINT-SJTU/Evo-temp.io

★ 1Forks 0PythonGitHub ↗Compare

README

Evo-RL

Open-source real-world RL toolkit for LeRobot101 (SO101), with continuous model, algorithm, and dataset releases.
LeRobot101 is live now. AgileX PiPER robot-arm support is coming soon.

Project Website • Reproduce Current Release • Community Program • WeChat Draft (ZH)

license python platform status

Positioning

Evo-RL is a continuous open-source program, not a one-time demo release.

  • Current platform: LeRobot101 / SO101.
  • Next platform: AgileX PiPER robot arms (coming soon, https://www.agilex.ai/).
  • Goal: build a community where tasks on both platforms can be uploaded, reproduced, and benchmarked with shared protocols.

Core Selling Points

  1. To our knowledge (within current open-source LeRobot ecosystem), Evo-RL is the first project continuously pushing real-world RL releases on LeRobot101.
  2. End-to-end loop already implemented: data collection -> value -> indicator -> policy -> real-world re-collection.
  3. Engineering-first design: intervention/outcome labels, dataset quality report, indicator-conditioned training integration, and reproducible CLI chain.
  4. Community-first roadmap: open task uploads and cross-platform benchmark evolution.

What Is Already Implemented

As of commit 852b23cb (2026-02-26), compared to main, core RL tooling includes:

  • 101 core files changed (website and outreach files excluded)
  • +6336 / -3224 lines
  • New CLIs:
    • lerobot-human-inloop-record
    • lerobot-dataset-report
    • lerobot-value-train
    • lerobot-value-infer
  • Indicator-conditioned policy path integrated into lerobot-train
  • Value/advantage/indicator write-back pipeline for iterative real-world training

Recompute stats with:

git diff --shortstat main..kye/main -- . \
  ':(exclude)website' \
  ':(exclude)docs/outreach' \
  ':(exclude).github/workflows/deploy-website-pages.yml' \
  ':(exclude)README.md'

Core Pipeline

[HIL collection on real robot]
         |
         v
[Dataset quality report]
         |
         v
[Value training]
         |
         v
[Value / Advantage / Indicator annotation]
         |
         v
[Indicator-conditioned policy training]
         |
         v
[Real-world rollout + next iteration]

Reproduce Current Release

Tested environment

  • Ubuntu 22.04
  • Python 3.10
  • CUDA 12.x class environment
  • SO-series leader/follower arm + USB camera (640x480@30)

1) Install

git clone https://github.com/Elvin-yk/evo-lerobot.git
cd evo-lerobot
conda activate lerobot
pip install -e .
pip install -e ".[pi]"      # value pipeline deps
pip install -e ".[feetech]" # SO101 related deps

2) Prepare variables

export DATASET_REPO_ID=<your_dataset_repo_id>
export DATASET_ROOT=~/.cache/huggingface/lerobot

3) Hardware preflight (highly recommended)

lerobot-find-port
lerobot-find-cameras

Confirm:

  • Correct follower/leader serial ports
  • Correct camera index/path
  • Current user has serial device permission

4) Human-in-loop recording

lerobot-human-inloop-record \
  --robot.type=so100_follower \
  --robot.port=/dev/ttyACM0 \
  --robot.cameras="{front: {type: opencv, index_or_path: 0, width: 640, height: 480, fps: 30}}" \
  --teleop.type=so100_leader \
  --teleop.port=/dev/ttyACM1 \
  --dataset.repo_id=${DATASET_REPO_ID} \
  --dataset.root=${DATASET_ROOT} \
  --dataset.single_task="Pick and place the red block" \
  --dataset.num_episodes=20 \
  --dataset.push_to_hub=false

Hotkeys:

  • i: toggle intervention
  • s: mark success and end episode
  • f: mark failure and end episode

5) Dataset quality report

lerobot-dataset-report --dataset ${DATASET_REPO_ID} --root ${DATASET_ROOT}
lerobot-dataset-report --dataset ${DATASET_REPO_ID} --root ${DATASET_ROOT} --json

Expected output:

  • Terminal report with success/failure/intervention metrics
  • JSON report for logging and experiment cards

6) Value training

lerobot-value-train \
  --dataset.repo_id=${DATASET_REPO_ID} \
  --dataset.root=${DATASET_ROOT} \
  --dataset.download_videos=true \
  --value.type=pistar06 \
  --batch_size=16 \
  --steps=2000 \
  --output_dir=outputs/value_train/evo_value_demo \
  --job_name=evo_value_demo

Expected output:

  • Value checkpoints under outputs/value_train/evo_value_demo

7) Value infer + advantage-indicator labels

lerobot-value-infer \
  --dataset.repo_id=${DATASET_REPO_ID} \
  --dataset.root=${DATASET_ROOT} \
  --inference.checkpoint_path=outputs/value_train/evo_value_demo \
  --inference.checkpoint_ref=last \
  --runtime.batch_size=64 \
  --acp.enable=true

Expected output:

  • Dataset fields written/updated:
    • complementary_info.value
    • complementary_info.advantage
    • complementary_info.acp_indicator

Optional visualization export:

lerobot-value-infer \
  --dataset.repo_id=${DATASET_REPO_ID} \
  --dataset.root=${DATASET_ROOT} \
  --inference.checkpoint_path=outputs/value_train/evo_value_demo \
  --inference.checkpoint_ref=last \
  --acp.enable=true \
  --viz.enable=true \
  --viz.episodes=all \
  --viz.output_dir=outputs/value_infer/viz

8) Indicator-conditioned policy training

CLI flags use --acp.*, where acp_indicator is the binary advantage-indicator field.

lerobot-train \
  --dataset.repo_id=${DATASET_REPO_ID} \
  --dataset.root=${DATASET_ROOT} \
  --policy.type=pi05 \
  --policy.pretrained_path=lerobot/pi05_base \
  --acp.enable=true \
  --acp.indicator_field=complementary_info.acp_indicator \
  --steps=3000

Current Evidence Snapshot

  • Platform: dual-arm SO101 task iteration
  • D0 dataset scale: 300 episodes, 413,134 frames, ~3.82h at 30 FPS
  • Internal observed trend: 100% data + Full FT forms a viable baseline in current setup

More benchmark cards and protocol details are published on the project website as releases progress.

Community Program

We are opening a community track for tasks on:

Current submission flow:

  1. Open an Issue with task setup and success definition.
  2. Share dataset link + exact training/inference commands.
  3. Submit result table + short video for benchmark integration.

Website

Project page is in website/:

cd website
python -m http.server 8000

Relationship to Upstream LeRobot

  • This repository is LeRobot-derived and keeps the lerobot-* CLI namespace.
  • Evo-RL extends upstream with real-world RL loop tooling and continuous release workflow.

Citation

@misc{evorl2026,
  title        = {Evo-RL: Continuous Open-Source Real-World RL on LeRobot101 and Beyond},
  author       = {Evo-RL Contributors},
  year         = {2026},
  howpublished = {\url{https://github.com/Elvin-yk/evo-lerobot}}
}

License

Apache-2.0. See LICENSE.

Contributors

alibertsCadenealexander-soareimstevenpmworkmichel-aractingipkooijElvin-ykAdilZouitinejadechoghariCarolinePascalfracapuanomishig25pre-commit-ci[bot]helper2424ben-znepyopemshukortc-huangthomwolfreeceomahoneyradekosmulskijackvialCharlesCNortons1lent4gntwut19sotanakamuraqgallouedecTavish9marinabardanaaubakirova

Issues