diff --git a/docs/source/libero.mdx b/docs/source/libero.mdx index b95af1d27..6dba5c62e 100644 --- a/docs/source/libero.mdx +++ b/docs/source/libero.mdx @@ -114,38 +114,58 @@ LIBERO supports two control modes — `relative` (default) and `absolute`. Diffe ### Recommended evaluation episodes -For reproducible benchmarking, use **10 episodes per task** across all four standard suites (Spatial, Object, Goal, Long). This gives 400 total episodes and matches the protocol used for published results. +For reproducible benchmarking, use **10 episodes per task** across all four standard suites (Spatial, Object, Goal, Long). This gives 400 total episodes and matches the protocol used for published results. Success rates may vary by a few percent across evaluation seeds, so we recommend averaging over 3 seeds. + + + To compare two policies on the same episodes, use the same `--seed`, keep + `--env.init_states=true`, and run each task in a single batch + (`--eval.batch_size` equal to episodes per task). + ## Training ### Dataset -We provide a preprocessed LIBERO dataset fully compatible with LeRobot: +Two preprocessed LIBERO datasets are fully compatible with LeRobot. They contain the same demonstrations with the same schema and differ in how camera frames are stored: -- [HuggingFaceVLA/libero](https://huggingface.co/datasets/HuggingFaceVLA/libero) +| | [lerobot/libero](https://huggingface.co/datasets/lerobot/libero) | [HuggingFaceVLA/libero](https://huggingface.co/datasets/HuggingFaceVLA/libero) | +| ------------------------- | ---------------------------------------------------------------- | ------------------------------------------------------------------------------ | +| episodes / frames / tasks | 1,693 / 273,465 / 40 | 1,693 / 273,465 / 40 | +| cameras | 2× 256×256×3 | 2× 256×256×3 | +| state / action dims | 8 / 7 | 8 / 7 | +| dataset format | v3.0 | v3.0 | +| camera encoding | MP4 video | PNG in parquet | +| download size | **1.9 GB** | 69.9 GB | +| extra dependency | video backend (`torchcodec` or `pyav`) | none | + +**We recommend [lerobot/libero](https://huggingface.co/datasets/lerobot/libero)**: **37× smaller download** with **equivalent loading speed** (~330 samples/s per worker). Video re-encoding is slightly lossy; use the image-based variant if you cannot install a video decoding backend. For reference, the original dataset published by Physical Intelligence: - [physical-intelligence/libero](https://huggingface.co/datasets/physical-intelligence/libero) + +Pin `--dataset.revision=` when reporting results — Hub datasets can be re-uploaded, and success rates are only comparable against the same data revision. + + ### Example training command +Train SmolVLA on the recommended dataset: + ```bash lerobot-train \ --policy.type=smolvla \ - --policy.repo_id=${HF_USER}/libero-test \ --policy.load_vlm_weights=true \ - --dataset.repo_id=HuggingFaceVLA/libero \ - --env.type=libero \ - --env.task=libero_10 \ - --output_dir=./outputs/ \ + --policy.push_to_hub=false \ + --dataset.repo_id=lerobot/libero \ + --dataset.video_backend=torchcodec \ + --output_dir=./outputs/libero_smolvla \ --steps=100000 \ - --batch_size=4 \ - --eval.batch_size=1 \ - --eval.n_episodes=1 \ - --env_eval_freq=1000 + --batch_size=64 ``` +To share the result on the Hub, replace `--policy.push_to_hub=false` with `--policy.repo_id=${HF_USER}/libero-smolvla`. Evaluate saved checkpoints with `lerobot-eval` as shown in the [Evaluation](#evaluation) section. + ## Reproducing published results We reproduce the results of Pi0.5 on the LIBERO benchmark. We take the Physical Intelligence LIBERO base model (`pi05_libero`) and finetune for an additional 6k steps in bfloat16, with batch size of 256 on 8 H100 GPUs using the [HuggingFace LIBERO dataset](https://huggingface.co/datasets/HuggingFaceVLA/libero).