Learning and Optimization for Robot Motion Planning and Control.
LearnOpt implements and explores optimization-based planning and learning-based control for quadrotors in MuJoCo. It provides reusable planning libraries, traditional and reinforcement-learning controllers, wind disturbances, and configurable training and deployment experiments.
Built on Mdhvince/UAV-Autonomous-control, the project extends the original simulation and tracking framework with FIRI, GCOPTER/MINCO, GCS, BMTP, nonlinear MPC, and PPO with MLP or differentiable MPC policies.
- Construct collision-free convex flight corridors with FIRI
- Optimize constrained multicopter trajectories with GCOPTER/MINCO
- Plan routes and Bézier trajectories using GCS
- Track trajectories with acados nonlinear MPC
- Simulate deterministic and randomized wind disturbances
- Train PPO trajectory-tracking policies with an MLP baseline
- Integrate differentiable MPC into PPO policies (ACMPC)
- Independently implement biconvex minimum-time planning (BMTP)
- Add a unified YAML entry for planning, training, deployment, and replay
- Extend RL training and deployment to tasks beyond trajectory tracking
- Implement diffusion-based motion planning
The current CLI supports trajectory tracking. A generic MuJoCo task interface is available for developing additional tasks. Implementation details and experiment workflows are documented in the planning library guide and contributor guide.
The videos show the implemented planning and tracking pipelines in MuJoCo. The orange curve is the planned trajectory and the blue curve is the actual flight path. Convex regions are not drawn in the videos; they are shown separately in the FIRI figures below.
gcopter_mpc.mp4
Planner: FIRI + GCOPTER
Controller: MPC
minisnap_cascaded.mp4
Planner: Minimum Snap baseline
Controller: Cascaded
gcs_mpc.mp4
Planner: GCS
Controller: Nonlinear MPC
FIRI builds collision-free convex regions around seed points or path segments. Each region has the half-space representation
It alternates obstacle-separating half-spaces with maximum-volume inscribed ellipsoid (MVIE) optimization. Overlapping regions form a safe corridor for trajectory optimization.
GCOPTER optimizes smooth multicopter trajectories inside the corridor using MINCO. Polynomial coefficients are recovered from compact spatial and temporal variables:
This implementation uses
L-BFGS optimizes the compact variables; the mission adapter samples the result into position, velocity, acceleration, and yaw references.
GCS starts from a cover of collision-free convex sets
This is a mixed-integer convex program. The implementation replaces
Actor-Critic Model Predictive Control (ACMPC) places a differentiable MPC layer at the end of the actor. A neural cost map turns the observation into time-varying quadratic cost residuals; together with known quadrotor dynamics and bounded inputs, they define the MPC action mean:
The first optimized input is the Gaussian actor mean, while a separate critic estimates the long-horizon return. PPO therefore learns the MPC cost map from reward rather than directly regressing an action. Gradients pass through the solver to the cost network:
This repository follows the paper's one-iLQR-update actor and uses an analytic differentiable backward pass. Its actor learns diagonal quadratic-cost residuals around a tracking cost; the critic is an MLP trained by PPO. Model-Predictive Value Expansion (MPVE) from the extended paper is not included.
BMTP jointly represents a Bézier trajectory
The term
Each update is convex; the overall alternating method is local and depends on its collision-free initialization.
Python 3.13, uv, and a graphical desktop for MuJoCo are required.
uv sync --python 3.13Set scene, planner, and controller in configs/flight.yaml, then run:
.venv/bin/python -m uav_ac.main --config configs/flight.yamlscene names an XML file under uav_ac/simulation/models/, with or without the .xml suffix. The main choices are:
| Field | Values | Notes |
|---|---|---|
planner |
mini_snap, gcopter, gcs, bmtp |
GCS needs scene guide regions; BMTP needs bmtp_route_* sites |
controller |
cascaded, mpc, rl |
MPC requires acados; RL requires rl.checkpoint |
wind |
none, fixed_gust |
Add wind_options only to override fixed-gust defaults |
visualize |
true, false |
Shows corridor geometry; BMTP always shows its dashed initialization and solid result |
Planner/controller-specific blocks can coexist as presets; only the block belonging to the selected planner or controller is used. configs/flight.yaml includes complete presets for gcopter, gcs, bmtp, cascaded, mpc, and rl. For example:
scene: bmtp_village
planner: bmtp
controller: cascaded
speed: 3.0
bmtp:
initial_route: 3
segments: 8
clearance: 0.15Use scene: gcs_building with planner: gcs. gcopter: and gcs: can override their native dataclass settings; mpc: configures the nonlinear controller. Traditional controllers default to a 0.01 s control_dt, which can be overridden with an integer multiple of the XML timestep. RL checkpoints own their control period. Viewer runs do not create output directories or overwrite previous results. Use uav_ac.record_experiments for videos.
For GCOPTER, keep speed as the shared velocity bound. gcopter: exposes trajectory scale (length_per_piece, time_weight), dynamic limits (max_acceleration, max_body_rate), soft-constraint weights, and optimizer convergence settings. gcs: exposes the Bézier graph optimization and solver settings; mpc: exposes NMPC horizon, tracking weights, and solver settings; cascaded: exposes response time constants, damping, and altitude integration. Mass, thrust, tilt, and flight-speed limits remain in the selected XML vehicle.
Build acados using its installation instructions, then install the Python interface into the project environment:
export ACADOS_SOURCE_DIR=/path/to/acados
export LD_LIBRARY_PATH="$ACADOS_SOURCE_DIR/lib:${LD_LIBRARY_PATH:-}"
uv pip install -e "$ACADOS_SOURCE_DIR/interfaces/acados_template"
.venv/bin/python -c "from acados_template import AcadosOcpSolver"Replace /path/to/acados with your installation path. Select the controller in configs/flight.yaml:
controller: mpc
mpc:
horizon_steps: 10
nlp_solver_type: SQP_RTITrajectory banks and checkpoints under runs/ are local artifacts and are not included in the repository.
First prepare a trajectory bank. The default configuration generates 200 training, 20 validation, and 20 test trajectories in the open-field scene and validates them with acados MPC:
.venv/bin/python -m uav_ac.rl.training --config configs/ppo_trajectory.yaml \
--run-dir runs/ppo_trajectory/multitraj01 --prepare-onlyGeneration can take substantial time. Train ACMPC with its independent training configuration and a new run directory:
.venv/bin/python -m uav_ac.rl.training \
--config configs/acmpc_trajectory.yaml \
--run-dir runs/acmpc_trajectory/multitraj01Edit configs/acmpc_trajectory.yaml for its trajectory bank, device, MPC, PPO, wind, and training scale. Use configs/ppo_trajectory.yaml for the MLP baseline.
To deploy either an MLP or ACMPC policy, let main plan the reference and select the explicit checkpoint:
scene: open_field
planner: gcopter
controller: rl
rl:
checkpoint: ../runs/acmpc_trajectory/multitraj01/best_model.zip
device: cudaCheckpoint paths are relative to the YAML file. Keep rl_config.json beside the model; the deployed controller adopts and validates the trained control period automatically. Dataset-based metrics, interactive evaluation, and recording remain available through uav_ac.rl.mlp_baseline.evaluate.
PYTEST_DISABLE_PLUGIN_AUTOLOAD=1 .venv/bin/python -m pytest -o addopts='' -qSee AGENT.md for contributor guidance, module contracts, and coverage checks.
configs/
├── flight.yaml Short interactive-flight configuration
├── ppo_trajectory.yaml MLP training and trajectory-bank settings
├── acmpc_trajectory.yaml ACMPC training settings
└── bmtp.yaml Standalone BMTP experiment settings
uav_ac/
├── main.py Scene, planner, controller, wind, and viewer entry
├── scenes/ XML loading and scene metadata
├── tasks/ Task protocol and trajectory tracking
├── envs/ Generic MuJoCo Gym environment
├── planning/
│ ├── geometry/ Convex geometry and collision utilities
│ ├── search/ RRT* path search
│ ├── corridor/firi/ Convex safe-corridor construction
│ ├── trajectory/ GCOPTER, GCS, BMTP, and Minimum Snap
│ └── pipeline/ Mission and trajectory conversion
├── control/ Cascaded/MPC control and tracking interfaces
├── rl/
│ ├── training.py Shared MLP/ACMPC trainer and bank preparation
│ ├── common/ Trajectory banks and initialization assets
│ ├── acmpc/ Differentiable MPC policy and solver
│ └── mlp_baseline/ Existing training/evaluation entry points
├── simulation/
│ └── models/ XML scenes and vehicle definitions
├── quadrotor/ Vehicle state and actuator allocation
└── visualization/ Planning overlays and plots
tests/ Unit and integration tests
docs/ Figures and documentation media
runs/ Local trajectory banks and trained models
For reusable APIs and extension points, see the planning guide and contributor guide.
-
Mdhvince, UAV-Autonomous-control. GitHub repository.
https://github.com/Mdhvince/UAV-Autonomous-control -
Z. Wang, X. Zhou, C. Xu, and F. Gao, “Geometrically Constrained Trajectory Optimization for Multicopters,” IEEE Transactions on Robotics, vol. 38, no. 5, pp. 3259–3278, 2022.
DOI: https://doi.org/10.1109/TRO.2022.3160022 -
Q. Wang, Z. Wang, M. Wang, J. Ji, Z. Han, T. Wu, R. Jin, Y. Gao, C. Xu, and F. Gao, “Fast Iterative Region Inflation for Computing Large 2-D/3-D Convex Regions of Obstacle-Free Space,” IEEE Transactions on Robotics, vol. 41, pp. 3223–3243, 2025.
DOI: https://doi.org/10.1109/TRO.2025.3562482 -
T. Marcucci, M. Petersen, D. von Wrangel, and R. Tedrake, “Motion Planning around Obstacles with Convex Optimization,” Science Robotics, vol. 8, no. 84, eadf7843, 2023.
DOI: https://doi.org/10.1126/scirobotics.adf7843 -
D. Mellinger and V. Kumar, “Minimum Snap Trajectory Generation and Control for Quadrotors,” in 2011 IEEE International Conference on Robotics and Automation (ICRA), pp. 2520–2525, 2011.
DOI: https://doi.org/10.1109/ICRA.2011.5980409 -
R. Verschueren, G. Frison, D. Kouzoupis, J. Frey, N. van Duijkeren, A. Zanelli, B. Novoselnik, T. Albin, R. Quirynen, and M. Diehl, “acados—a modular open-source framework for fast embedded optimal control,” Mathematical Programming Computation, vol. 14, no. 1, pp. 147–183, 2022.
DOI: https://doi.org/10.1007/s12532-021-00208-8 -
P. Werner, T. Marcucci, and D. Rus, “Biconvex Optimization for Smooth Minimum-Time Trajectories around Convex Obstacles,” arXiv preprint arXiv:2608.02834, 2026. https://arxiv.org/abs/2608.02834
-
A. Romero, E. Aljalbout, Y. Song, and D. Scaramuzza, “Actor–Critic Model Predictive Control: Differentiable Optimization Meets Reinforcement Learning for Agile Flight,” IEEE Transactions on Robotics, 2025. DOI: https://doi.org/10.1109/TRO.2025.3644945
