mirror of
https://github.com/huggingface/lerobot.git
synced 2026-08-08 17:39:44 +00:00
266be2bd17
* feat(train): add opt-in EMA of the policy weights (--ema.enable=true) Maintain an EMA shadow via diffusers' EMAModel (lazy import, no new dependency) with the reference Diffusion Policy schedule. Saves the shadow for exact resume plus a loadable pretrained_model_ema/ per checkpoint, evaluates the EMA weights during env eval, and pushes them to a sibling <repo_id>-ema repo. Fixes huggingface/lerobot#4259. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(diffusion): document the --ema.enable training flag Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(tests): skip EMA training tests when accelerate/diffusers are missing * feat(train): support constant EMA decay (--ema.decay) for openpi-style policies * fix(train): gate EMA step on sync_gradients; use parallel_dims.is_sharded guard --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
108 lines
5.3 KiB
Markdown
108 lines
5.3 KiB
Markdown
# π₀.₅ (pi05)
|
|
|
|
This repository contains the Hugging Face port of **π₀.₅**, adapted from [OpenPI](https://github.com/Physical-Intelligence/openpi) by the Physical Intelligence.
|
|
It is designed as a **Vision-Language-Action model with open-world generalization**.
|
|
|
|
---
|
|
|
|
## Model Overview
|
|
|
|
| Feature | π₀ | π₀.₅ |
|
|
| -------------------- | ------------------------------------------------------ | ----------------------------------------- |
|
|
| Time Conditioning | Concatenates time with actions via `action_time_mlp_*` | Uses `time_mlp_*` for AdaRMS conditioning |
|
|
| AdaRMS | Not used | Used in action expert |
|
|
| Tokenizer Length | 48 tokens | 200 tokens |
|
|
| Discrete State Input | False (Uses `state_proj` layer) | True |
|
|
| Parameter Count | Higher (includes state embedding) | Lower (no state embedding) |
|
|
|
|
---
|
|
|
|
## Relative Actions
|
|
|
|
π₀.₅ supports training with **relative actions**, where the model learns relative offsets
|
|
from the current robot state instead of absolute joint positions. This mirrors the
|
|
relative-action transform in OpenPI (`DeltaActions`) and can improve performance.
|
|
|
|
### How it works
|
|
|
|
1. **During preprocessing**, absolute actions are converted to relative offsets:
|
|
`relative = action - state` (for selected joints).
|
|
2. The relative actions are normalized using statistics computed from the relative distribution.
|
|
3. **During postprocessing**, predicted relative actions are converted back to absolute:
|
|
`absolute = relative + state`.
|
|
|
|
Joints listed in `relative_exclude_joints` (e.g., gripper) are kept absolute.
|
|
|
|
### Configuration
|
|
|
|
| Parameter | Type | Default | Description |
|
|
| ------------------------- | ----------- | ------------- | ---------------------------------------------------------------- |
|
|
| `use_relative_actions` | `bool` | `False` | Enable relative-action training |
|
|
| `relative_exclude_joints` | `list[str]` | `["gripper"]` | Joint names to keep absolute (matched by substring) |
|
|
| `action_feature_names` | `list[str]` | `None` | Auto-populated from dataset metadata at runtime by `make_policy` |
|
|
|
|
### Training example
|
|
|
|
```bash
|
|
python -m lerobot.scripts.lerobot_train \
|
|
--policy.type=pi05 \
|
|
--dataset.repo_id=your_org/your_dataset \
|
|
--policy.use_relative_actions=true \
|
|
--policy.relative_exclude_joints='["gripper"]'
|
|
```
|
|
|
|
When `use_relative_actions=true`, the training script automatically:
|
|
|
|
- Computes relative action statistics from the dataset (sampled chunk-level relative actions)
|
|
- Replaces the standard action stats with relative stats for normalization
|
|
- Broadcasts these stats across all ranks in distributed training
|
|
|
|
---
|
|
|
|
## EMA of the policy weights
|
|
|
|
OpenPI maintains an exponential moving average of the weights during training (`ema_decay=0.99` by default) and keeps the EMA copy for inference. To reproduce this with the LeRobot trainer, enable the EMA shadow with a constant decay:
|
|
|
|
```bash
|
|
python -m lerobot.scripts.lerobot_train \
|
|
--policy.type=pi05 \
|
|
--dataset.repo_id=your_org/your_dataset \
|
|
--ema.enable=true \
|
|
--ema.decay=0.99
|
|
```
|
|
|
|
Checkpoints then contain a directly loadable copy of the EMA weights in `pretrained_model_ema/` next to the live ones. Note that the shadow is a full extra copy of the parameters on the GPU. Like OpenPI (which disables EMA in its LoRA configs), EMA is not supported together with PEFT adapters.
|
|
|
|
---
|
|
|
|
## Citation
|
|
|
|
If you use this work, please cite both **OpenPI** and the π₀.₅ paper:
|
|
|
|
```bibtex
|
|
@misc{openpi2024,
|
|
author = {Physical Intelligence Lab},
|
|
title = {OpenPI: PyTorch Implementation of π0 and π0.5 Policies},
|
|
year = {2024},
|
|
publisher = {GitHub},
|
|
howpublished = {\url{https://github.com/Physical-Intelligence/openpi}},
|
|
license = {Apache-2.0}
|
|
}
|
|
|
|
@misc{intelligence2025pi05visionlanguageactionmodelopenworld,
|
|
title = {π₀.₅: a Vision-Language-Action Model with Open-World Generalization},
|
|
author = {Physical Intelligence and Kevin Black and Noah Brown and James Darpinian and Karan Dhabalia and Danny Driess and Adnan Esmail and Michael Equi and Chelsea Finn and Niccolo Fusai and Manuel Y. Galliker and Dibya Ghosh and Lachy Groom and Karol Hausman and Brian Ichter and Szymon Jakubczak and Tim Jones and Liyiming Ke and Devin LeBlanc and Sergey Levine and Adrian Li-Bell and Mohith Mothukuri and Suraj Nair and Karl Pertsch and Allen Z. Ren and Lucy Xiaoyang Shi and Laura Smith and Jost Tobias Springenberg and Kyle Stachowicz and James Tanner and Quan Vuong and Homer Walke and Anna Walling and Haohuan Wang and Lili Yu and Ury Zhilinsky},
|
|
year = {2025},
|
|
eprint = {2504.16054},
|
|
archivePrefix= {arXiv},
|
|
primaryClass = {cs.LG},
|
|
url = {https://arxiv.org/abs/2504.16054},
|
|
}
|
|
```
|
|
|
|
---
|
|
|
|
## License
|
|
|
|
This port follows the **Apache 2.0 License**, consistent with the original [OpenPI repository](https://github.com/Physical-Intelligence/openpi).
|