mirror of
https://github.com/huggingface/lerobot.git
synced 2026-07-26 11:16:00 +00:00
docs(robocerebra): align page with adding_benchmarks template
Rework docs/source/robocerebra.mdx to follow the standard benchmark doc structure: intro + links + available tasks + installation + eval + recommended episodes + policy I/O + training + reproducing results. - Point everything at lerobot/smolvla_robocerebra (the released checkpoint), not the personal pepijn223 mirror. - Add the --env.fps=20 and --env.obs_type=pixels_agent_pos flags that CI actually uses, so copy-paste eval reproduces CI. - Split the "Training" block out of the recipe section into its own section with the feature table. - Add an explicit "Reproducing published results" section pointing at the CI smoke eval. Made-with: Cursor
This commit is contained in:
@@ -1,14 +1,25 @@
|
|||||||
# RoboCerebra
|
# RoboCerebra
|
||||||
|
|
||||||
[RoboCerebra](https://robocerebra-project.github.io/) is a long-horizon manipulation benchmark for evaluating high-level reasoning, planning, and memory in VLAs. Episodes chain multiple sub-goals with language-grounded intermediate instructions, building on top of LIBERO's simulator stack.
|
[RoboCerebra](https://robocerebra-project.github.io/) is a long-horizon manipulation benchmark that evaluates **high-level reasoning, planning, and memory** in VLAs. Episodes chain multiple sub-goals with language-grounded intermediate instructions, built on top of LIBERO's simulator stack (MuJoCo + robosuite, Franka Panda 7-DOF).
|
||||||
|
|
||||||
- Paper: [RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation](https://arxiv.org/abs/2506.06677)
|
- Paper: [RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation](https://arxiv.org/abs/2506.06677)
|
||||||
- Project website: [robocerebra-project.github.io](https://robocerebra-project.github.io/)
|
- Project website: [robocerebra-project.github.io](https://robocerebra-project.github.io/)
|
||||||
- Dataset: [`lerobot/robocerebra_unified`](https://huggingface.co/datasets/lerobot/robocerebra_unified) — LeRobot v3.0, 6,660 episodes (571,116 frames) at 20 fps, 1,728 language-grounded sub-tasks
|
- Dataset: [`lerobot/robocerebra_unified`](https://huggingface.co/datasets/lerobot/robocerebra_unified) — LeRobot v3.0, 6,660 episodes / 571,116 frames at 20 fps, 1,728 language-grounded sub-tasks.
|
||||||
|
- Pretrained policy: [`lerobot/smolvla_robocerebra`](https://huggingface.co/lerobot/smolvla_robocerebra)
|
||||||
|
|
||||||
|
## Available tasks
|
||||||
|
|
||||||
|
RoboCerebra reuses LIBERO's simulator, so evaluation runs against the LIBERO `libero_10` long-horizon suite:
|
||||||
|
|
||||||
|
| Suite | CLI name | Tasks | Description |
|
||||||
|
| --------- | ----------- | ----- | ------------------------------------------------------------- |
|
||||||
|
| LIBERO-10 | `libero_10` | 10 | Long-horizon kitchen/living room tasks chaining 3–6 sub-goals |
|
||||||
|
|
||||||
|
Each RoboCerebra episode in the dataset is segmented into multiple sub-tasks with natural-language instructions, which the unified dataset exposes as independent supervision signals.
|
||||||
|
|
||||||
## Installation
|
## Installation
|
||||||
|
|
||||||
RoboCerebra reuses LIBERO's simulator stack, so install the `libero` extra:
|
RoboCerebra piggybacks on LIBERO, so the `libero` extra is all you need:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
pip install -e ".[libero]"
|
pip install -e ".[libero]"
|
||||||
@@ -18,29 +29,47 @@ pip install -e ".[libero]"
|
|||||||
RoboCerebra requires Linux (MuJoCo / robosuite). Set the rendering backend before training or evaluation:
|
RoboCerebra requires Linux (MuJoCo / robosuite). Set the rendering backend before training or evaluation:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
export MUJOCO_GL=egl # for headless servers
|
export MUJOCO_GL=egl # for headless servers (HPC, cloud)
|
||||||
```
|
```
|
||||||
|
|
||||||
</Tip>
|
</Tip>
|
||||||
|
|
||||||
## Evaluation
|
## Evaluation
|
||||||
|
|
||||||
RoboCerebra eval runs against the LIBERO `libero_10` long-horizon suite with RoboCerebra's camera naming (`image` + `wrist_image`) and an extra empty-camera slot for policies expecting three views:
|
RoboCerebra eval runs against LIBERO's `libero_10` suite with RoboCerebra's camera naming (`image` + `wrist_image`) and an extra empty-camera slot for policies trained with three views:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
lerobot-eval \
|
lerobot-eval \
|
||||||
--policy.path=pepijn223/smolvla_robocerebra \
|
--policy.path=lerobot/smolvla_robocerebra \
|
||||||
--env.type=libero \
|
--env.type=libero \
|
||||||
--env.task=libero_10 \
|
--env.task=libero_10 \
|
||||||
|
--env.fps=20 \
|
||||||
|
--env.obs_type=pixels_agent_pos \
|
||||||
--eval.batch_size=1 \
|
--eval.batch_size=1 \
|
||||||
--eval.n_episodes=10 \
|
--eval.n_episodes=10 \
|
||||||
'--rename_map={"observation.images.image": "observation.images.camera1", "observation.images.wrist_image": "observation.images.camera2"}' \
|
'--rename_map={"observation.images.image": "observation.images.camera1", "observation.images.wrist_image": "observation.images.camera2"}' \
|
||||||
--policy.empty_cameras=3
|
--policy.empty_cameras=3
|
||||||
```
|
```
|
||||||
|
|
||||||
|
### Recommended evaluation episodes
|
||||||
|
|
||||||
|
**10 episodes per task** across the `libero_10` suite (100 total) for reproducible benchmarking. Matches the protocol used in the RoboCerebra paper.
|
||||||
|
|
||||||
|
## Policy inputs and outputs
|
||||||
|
|
||||||
|
**Observations:**
|
||||||
|
|
||||||
|
- `observation.state` — 8-dim proprioceptive state (7 joint positions + gripper)
|
||||||
|
- `observation.images.image` — third-person view, 256×256 HWC uint8
|
||||||
|
- `observation.images.wrist_image` — wrist-mounted camera view, 256×256 HWC uint8
|
||||||
|
|
||||||
|
**Actions:**
|
||||||
|
|
||||||
|
- Continuous control in `Box(-1, 1, shape=(7,))` — end-effector delta (6D) + gripper (1D)
|
||||||
|
|
||||||
## Training
|
## Training
|
||||||
|
|
||||||
The unified dataset at [`lerobot/robocerebra_unified`](https://huggingface.co/datasets/lerobot/robocerebra_unified) exposes two RGB streams:
|
The unified dataset at [`lerobot/robocerebra_unified`](https://huggingface.co/datasets/lerobot/robocerebra_unified) exposes two RGB streams and language-grounded sub-task annotations:
|
||||||
|
|
||||||
| Feature | Shape | Description |
|
| Feature | Shape | Description |
|
||||||
| -------------------------------- | ------------- | -------------------- |
|
| -------------------------------- | ------------- | -------------------- |
|
||||||
@@ -59,3 +88,7 @@ lerobot-train \
|
|||||||
--env.task=libero_10 \
|
--env.task=libero_10 \
|
||||||
--output_dir=outputs/smolvla_robocerebra
|
--output_dir=outputs/smolvla_robocerebra
|
||||||
```
|
```
|
||||||
|
|
||||||
|
## Reproducing published results
|
||||||
|
|
||||||
|
The released checkpoint [`lerobot/smolvla_robocerebra`](https://huggingface.co/lerobot/smolvla_robocerebra) was trained on `lerobot/robocerebra_unified` and evaluated with the command in the [Evaluation](#evaluation) section. CI runs the same command with `--eval.n_episodes=1` as a smoke test on every PR touching the benchmark.
|
||||||
|
|||||||
Reference in New Issue
Block a user