From 1dd5711809b5daddb515cf92adbe7e63291aa7a3 Mon Sep 17 00:00:00 2001 From: Pepijn Date: Fri, 17 Apr 2026 13:45:01 +0100 Subject: [PATCH] docs(robocerebra): align page with adding_benchmarks template Rework docs/source/robocerebra.mdx to follow the standard benchmark doc structure: intro + links + available tasks + installation + eval + recommended episodes + policy I/O + training + reproducing results. - Point everything at lerobot/smolvla_robocerebra (the released checkpoint), not the personal pepijn223 mirror. - Add the --env.fps=20 and --env.obs_type=pixels_agent_pos flags that CI actually uses, so copy-paste eval reproduces CI. - Split the "Training" block out of the recipe section into its own section with the feature table. - Add an explicit "Reproducing published results" section pointing at the CI smoke eval. Made-with: Cursor --- docs/source/robocerebra.mdx | 49 +++++++++++++++++++++++++++++++------ 1 file changed, 41 insertions(+), 8 deletions(-) diff --git a/docs/source/robocerebra.mdx b/docs/source/robocerebra.mdx index 08007c740..015db01f4 100644 --- a/docs/source/robocerebra.mdx +++ b/docs/source/robocerebra.mdx @@ -1,46 +1,75 @@ # RoboCerebra -[RoboCerebra](https://robocerebra-project.github.io/) is a long-horizon manipulation benchmark for evaluating high-level reasoning, planning, and memory in VLAs. Episodes chain multiple sub-goals with language-grounded intermediate instructions, building on top of LIBERO's simulator stack. +[RoboCerebra](https://robocerebra-project.github.io/) is a long-horizon manipulation benchmark that evaluates **high-level reasoning, planning, and memory** in VLAs. Episodes chain multiple sub-goals with language-grounded intermediate instructions, built on top of LIBERO's simulator stack (MuJoCo + robosuite, Franka Panda 7-DOF). - Paper: [RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation](https://arxiv.org/abs/2506.06677) - Project website: [robocerebra-project.github.io](https://robocerebra-project.github.io/) -- Dataset: [`lerobot/robocerebra_unified`](https://huggingface.co/datasets/lerobot/robocerebra_unified) — LeRobot v3.0, 6,660 episodes (571,116 frames) at 20 fps, 1,728 language-grounded sub-tasks +- Dataset: [`lerobot/robocerebra_unified`](https://huggingface.co/datasets/lerobot/robocerebra_unified) — LeRobot v3.0, 6,660 episodes / 571,116 frames at 20 fps, 1,728 language-grounded sub-tasks. +- Pretrained policy: [`lerobot/smolvla_robocerebra`](https://huggingface.co/lerobot/smolvla_robocerebra) + +## Available tasks + +RoboCerebra reuses LIBERO's simulator, so evaluation runs against the LIBERO `libero_10` long-horizon suite: + +| Suite | CLI name | Tasks | Description | +| --------- | ----------- | ----- | ------------------------------------------------------------- | +| LIBERO-10 | `libero_10` | 10 | Long-horizon kitchen/living room tasks chaining 3–6 sub-goals | + +Each RoboCerebra episode in the dataset is segmented into multiple sub-tasks with natural-language instructions, which the unified dataset exposes as independent supervision signals. ## Installation -RoboCerebra reuses LIBERO's simulator stack, so install the `libero` extra: +RoboCerebra piggybacks on LIBERO, so the `libero` extra is all you need: ```bash pip install -e ".[libero]" ``` -RoboCerebra requires Linux (MuJoCo/robosuite). Set the rendering backend before training or evaluation: +RoboCerebra requires Linux (MuJoCo / robosuite). Set the rendering backend before training or evaluation: ```bash -export MUJOCO_GL=egl # for headless servers +export MUJOCO_GL=egl # for headless servers (HPC, cloud) ``` ## Evaluation -RoboCerebra eval runs against the LIBERO `libero_10` long-horizon suite with RoboCerebra's camera naming (`image` + `wrist_image`) and an extra empty-camera slot for policies expecting three views: +RoboCerebra eval runs against LIBERO's `libero_10` suite with RoboCerebra's camera naming (`image` + `wrist_image`) and an extra empty-camera slot for policies trained with three views: ```bash lerobot-eval \ - --policy.path=pepijn223/smolvla_robocerebra \ + --policy.path=lerobot/smolvla_robocerebra \ --env.type=libero \ --env.task=libero_10 \ + --env.fps=20 \ + --env.obs_type=pixels_agent_pos \ --eval.batch_size=1 \ --eval.n_episodes=10 \ '--rename_map={"observation.images.image": "observation.images.camera1", "observation.images.wrist_image": "observation.images.camera2"}' \ --policy.empty_cameras=3 ``` +### Recommended evaluation episodes + +**10 episodes per task** across the `libero_10` suite (100 total) for reproducible benchmarking. Matches the protocol used in the RoboCerebra paper. + +## Policy inputs and outputs + +**Observations:** + +- `observation.state` — 8-dim proprioceptive state (7 joint positions + gripper) +- `observation.images.image` — third-person view, 256×256 HWC uint8 +- `observation.images.wrist_image` — wrist-mounted camera view, 256×256 HWC uint8 + +**Actions:** + +- Continuous control in `Box(-1, 1, shape=(7,))` — end-effector delta (6D) + gripper (1D) + ## Training -The unified dataset at [`lerobot/robocerebra_unified`](https://huggingface.co/datasets/lerobot/robocerebra_unified) exposes two RGB streams: +The unified dataset at [`lerobot/robocerebra_unified`](https://huggingface.co/datasets/lerobot/robocerebra_unified) exposes two RGB streams and language-grounded sub-task annotations: | Feature | Shape | Description | | -------------------------------- | ------------- | -------------------- | @@ -59,3 +88,7 @@ lerobot-train \ --env.task=libero_10 \ --output_dir=outputs/smolvla_robocerebra ``` + +## Reproducing published results + +The released checkpoint [`lerobot/smolvla_robocerebra`](https://huggingface.co/lerobot/smolvla_robocerebra) was trained on `lerobot/robocerebra_unified` and evaluated with the command in the [Evaluation](#evaluation) section. CI runs the same command with `--eval.n_episodes=1` as a smoke test on every PR touching the benchmark.