docs(omx): adding some examples and scripts (#3566 )

* docs(omx): adding some examples and scripts * cleaning up and reviewing the cli args * adding __init__.py to example folder, adjusting the examples * adding reference to pretrained act policy * moving `.send_action` before `dataset.add_frame` for consistency Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> Signed-off-by: Maxime Ellerbach <maxime@ellerbach.net> * adjusting docstring Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> Signed-off-by: Maxime Ellerbach <maxime@ellerbach.net> * adressing hardcoded dataset fps * removed init as it worked without --------- Signed-off-by: Maxime Ellerbach <maxime@ellerbach.net>
docs: add policy & compute guide (#3534 )
2026-05-11 14:49:43 +00:00 · 2026-05-11 15:36:32 +02:00 · 2026-05-11 15:19:12 +02:00 · 2026-05-10 17:30:37 +02:00 · 2026-05-10 13:05:35 +02:00 · 2026-05-10 12:24:22 +02:00
24 changed files with 1826 additions and 648 deletions
@@ -382,6 +382,7 @@ jobs:
                --policy.path=\"\$ROBOTWIN_POLICY\" \
                --env.type=robotwin \
                --env.task=\"\$ROBOTWIN_TASKS\" \
+                --env.max_parallel_tasks=5 \
                --eval.batch_size=1 \
                --eval.n_episodes=1 \
                --eval.use_async_envs=false \
@@ -482,6 +483,7 @@ jobs:
                --policy.path=lerobot/smolvla_robocasa \
                --env.type=robocasa \
                --env.task=CloseFridge,OpenCabinet,OpenDrawer,TurnOnMicrowave,TurnOffStove,CloseToasterOvenDoor,SlideDishwasherRack,TurnOnSinkFaucet,NavigateKitchen,TurnOnElectricKettle \
+                --env.max_parallel_tasks=5 \
                --eval.batch_size=1 \
                --eval.n_episodes=1 \
                --eval.use_async_envs=false \
@@ -693,6 +695,7 @@ jobs:
                --env.task=\"\$ROBOMME_TASKS\" \
                --env.dataset_split=test \
                --env.task_ids=[0] \
+                --env.max_parallel_tasks=5 \
                --eval.batch_size=1 \
                --eval.n_episodes=1 \
                --eval.use_async_envs=false \
@@ -800,6 +803,7 @@ jobs:
                --env.type=libero_plus \
                --env.task=\"\$LIBERO_PLUS_SUITE\" \
                --env.task_ids=\"\$LIBERO_PLUS_TASK_IDS\" \
+                --env.max_parallel_tasks=5 \
                --eval.batch_size=1 \
                --eval.n_episodes=1 \
                --eval.use_async_envs=false \
@@ -900,6 +904,8 @@ jobs:
                --policy.path=lerobot/smolvla_vlabench \
                --env.type=vlabench \
                --env.task=select_fruit,select_toy,select_book,select_painting,select_drink,select_ingredient,select_billiards,select_poker,add_condiment,insert_flower \
+                --env.episode_length=50 \
+                --env.max_parallel_tasks=5 \
                --eval.batch_size=1 \
                --eval.n_episodes=1 \
                --eval.use_async_envs=false \
@@ -19,8 +19,8 @@ on:
  workflow_dispatch:

  # Runs at 02:00
-  schedule:
-    - cron: "0 2 * * *"
+  # schedule:
+  #   - cron: "0 2 * * *"

 env:
  CLOSE_ISSUE_MESSAGE: >
@@ -232,6 +232,8 @@ Match the policy to the user's **GPU memory** and **time budget**. Numbers below

 All policies typically train for **5–10 epochs** (see §7).

+> **Human-facing version:** the [Compute Hardware Guide](./docs/source/hardware_guide.mdx) reuses the table below and adds a cloud-GPU tier guide and a Hugging Face Jobs pointer.
+
 | Policy      | Batch | Update (ms) | Peak GPU mem (GB) | Best for                                                                                         |
 | ----------- | ----: | ----------: | ----------------: | ------------------------------------------------------------------------------------------------ |
 | `act`       |     4 |    **83.9** |          **0.94** | First-time users, laptops, single-task. Fast and reliable.                                       |
@@ -109,7 +109,7 @@ lerobot-train \

 Similarly to the hardware, you can easily implement your own policy & leverage LeRobot's data collection, training, and visualization tools, and share your model to the HF Hub

-For detailed policy setup guides, see the [Policy Documentation](https://huggingface.co/docs/lerobot/bring_your_own_policies).
+For detailed policy setup guides, see the [Policy Documentation](https://huggingface.co/docs/lerobot/bring_your_own_policies). For GPU/RAM requirements and expected training time per policy, see the [Compute Hardware Guide](https://huggingface.co/docs/lerobot/hardware_guide).

 ## Inference & Evaluation

@@ -35,7 +35,7 @@ USER root
 ARG ROBOTWIN_SHA=0aeea2d669c0f8516f4d5785f0aa33ba812c14b4
 RUN apt-get update \
    && apt-get install -y --no-install-recommends \
-         cuda-nvcc-12-4 cuda-cudart-dev-12-4 \
+         cuda-nvcc-12-6 cuda-cudart-dev-12-6 \
         libvulkan1 vulkan-tools \
    && mkdir -p /usr/share/vulkan/icd.d \
    && echo '{"file_format_version":"1.0.0","ICD":{"library_path":"libGLX_nvidia.so.0","api_version":"1.3.0"}}' \
@@ -8,7 +8,7 @@
  - local: il_robots
    title: Imitation Learning for Robots
  - local: bring_your_own_policies
-    title: Bring Your Own Policies
+    title: Adding a Policy
  - local: integrate_hardware
    title: Bring Your Own Hardware
  - local: hilserl
@@ -24,6 +24,12 @@
  - local: rename_map
    title: Using Rename Map and Empty Cameras
  title: "Tutorials"
+- sections:
+  - local: hardware_guide
+    title: Compute Hardware Guide
+  - local: torch_accelerators
+    title: PyTorch accelerators
+  title: "Compute & Hardware"
 - sections:
  - local: lerobot-dataset-v3
    title: Using LeRobotDataset
@@ -142,10 +148,6 @@
  - local: cameras
    title: Cameras
  title: "Sensors"
- sections:
-  - local: torch_accelerators
-    title: PyTorch accelerators
-  title: "Supported Hardware"
 - sections:
  - local: notebooks
    title: Notebooks
@@ -1,60 +1,37 @@
-# Bring Your Own Policies
+# Adding a Policy

-This tutorial explains how to integrate your own custom policy implementations into the LeRobot ecosystem, allowing you to leverage all LeRobot tools for training, evaluation, and deployment while using your own algorithms.
+This guide walks you through implementing a custom policy and getting it to work with LeRobot's training, evaluation, and deployment tools. There are two paths:

-## Step 1: Create a Policy Package
+- **Plugin (out-of-tree)** — ship your policy as a standalone `lerobot_policy_*` package. Faster, no PR required, easy to iterate. Right for experimentation, internal use, or when you want to publish independently.
+- **In-tree (contributed to LeRobot)** — land your policy directly in `src/lerobot/policies/`. Requires a PR, but makes your policy a first-class citizen of the library.

-Your custom policy should be organized as an installable Python package following LeRobot's plugin conventions.
+The plugin route is usually the right starting point — promote to in-tree once the policy has stabilized and there's clear value in shipping it with the library.

-### Package Structure
+Either way, the building blocks are the same: a configuration class, a policy class, and a processor factory. The first half of this guide covers those shared pieces; the second half covers the path-specific scaffolding ([Path A](#path-a-out-of-tree-plugin), [Path B](#path-b-contributing-in-tree)).

-Create a package with the prefix `lerobot_policy_` (IMPORTANT!) followed by your policy name:
+A note on tone: robot-learning is an actively evolving field, and "what a policy looks like" can shift with each new architecture. The conventions described here exist because they let `lerobot-train` and `lerobot-eval` work uniformly across very different models. When a new policy genuinely doesn't fit them, raise it (in your PR, or an issue) — the conventions are not sacred.

-```bash
-lerobot_policy_my_custom_policy/
-├── pyproject.toml
-└── src/
-    └── lerobot_policy_my_custom_policy/
-        ├── __init__.py
-        ├── configuration_my_custom_policy.py
-        ├── modeling_my_custom_policy.py
-        └── processor_my_custom_policy.py
-```
+---

-### Package Configuration
+## Anatomy of a policy

-Set up your `pyproject.toml`:
+Three building blocks make up every policy. The names below use `my_policy` as a placeholder — replace with your policy's name. That name is load-bearing: it must match the string you pass to `@PreTrainedConfig.register_subclass`, the `MyPolicy.name` class attribute, and the `make_<name>_pre_post_processors` factory function (more on each below).

-```toml
-[project]
-name = "lerobot_policy_my_custom_policy"
-version = "0.1.0"
-dependencies = [
-    # your policy-specific dependencies
-]
-requires-python = ">= 3.12"
+### Configuration class

-[build-system]
-build-backend = # your-build-backend
-requires = # your-build-system
-```
-
-## Step 2: Define the Policy Configuration
-
-Create a configuration class that inherits from [`PreTrainedConfig`](https://github.com/huggingface/lerobot/blob/main/src/lerobot/configs/policies.py) and registers your policy type:
-Here is a template to get you started, customize the parameters and methods as needed for your policy's architecture and training requirements.
+Inherit from [`PreTrainedConfig`](https://github.com/huggingface/lerobot/blob/main/src/lerobot/configs/policies.py) and register your policy type. Here is a template — customize the parameters and methods as needed for your policy's architecture and training requirements.

 ```python
-# configuration_my_custom_policy.py
+# configuration_my_policy.py
 from dataclasses import dataclass, field
 from lerobot.configs import PreTrainedConfig
 from lerobot.optim import AdamWConfig
 from lerobot.optim import CosineDecayWithWarmupSchedulerConfig

-@PreTrainedConfig.register_subclass("my_custom_policy")
+@PreTrainedConfig.register_subclass("my_policy")
@dataclass
-class MyCustomPolicyConfig(PreTrainedConfig):
-    """Configuration class for MyCustomPolicy.
+class MyPolicyConfig(PreTrainedConfig):
+    """Configuration class for MyPolicy.

    Args:
        n_obs_steps: Number of observation steps to use as input
@@ -77,16 +54,20 @@ class MyCustomPolicyConfig(PreTrainedConfig):
            raise ValueError("n_action_steps cannot exceed horizon")

    def validate_features(self) -> None:
-        """Validate input/output feature compatibility."""
+        """Validate input/output feature compatibility.
+
+        Call this explicitly from your policy's __init__ — the base class does not.
+        """
        if not self.image_features:
-            raise ValueError("MyCustomPolicy requires at least one image feature.")
+            raise ValueError("MyPolicy requires at least one image feature.")
        if self.action_feature is None:
-            raise ValueError("MyCustomPolicy requires 'action' in output_features.")
+            raise ValueError("MyPolicy requires 'action' in output_features.")

    def get_optimizer_preset(self) -> AdamWConfig:
        return AdamWConfig(lr=self.optimizer_lr, weight_decay=self.optimizer_weight_decay)

    def get_scheduler_preset(self):
+        """Return a LRSchedulerConfig from lerobot.optim, or None."""
        return None

    @property
@@ -101,8 +82,7 @@ class MyCustomPolicyConfig(PreTrainedConfig):

    @property
    def action_delta_indices(self) -> list[int]:
-        """Relative timestep offsets for the action chunk the dataset loader returns.
-        """
+        """Relative timestep offsets for the action chunk the dataset loader returns."""
        return list(range(self.horizon))

    @property
@@ -110,32 +90,34 @@ class MyCustomPolicyConfig(PreTrainedConfig):
        return None
 ```

-## Step 3: Implement the Policy Class
+The string you pass to `@register_subclass` must match `MyPolicy.name` (next section) and is what users supply as `--policy.type` on the CLI. Default to `AdamW` from `lerobot.optim` for `get_optimizer_preset` unless you genuinely need otherwise.

-Create your policy implementation by inheriting from [`PreTrainedPolicy`](https://github.com/huggingface/lerobot/blob/main/src/lerobot/policies/pretrained.py):
+### Policy class
+
+Inherit from [`PreTrainedPolicy`](https://github.com/huggingface/lerobot/blob/main/src/lerobot/policies/pretrained.py) and set two class attributes — both are checked by `__init_subclass__`:

 ```python
-# modeling_my_custom_policy.py
+# modeling_my_policy.py
 import torch
 import torch.nn as nn
 from typing import Any

 from lerobot.policies import PreTrainedPolicy
 from lerobot.utils.constants import ACTION
-from .configuration_my_custom_policy import MyCustomPolicyConfig
+from .configuration_my_policy import MyPolicyConfig

-class MyCustomPolicy(PreTrainedPolicy):
-    config_class = MyCustomPolicyConfig  # must match the string in @register_subclass
-    name = "my_custom_policy"
+class MyPolicy(PreTrainedPolicy):
+    config_class = MyPolicyConfig  # must match the string in @register_subclass
+    name = "my_policy"

-    def __init__(self, config: MyCustomPolicyConfig, dataset_stats: dict[str, Any] = None):
+    def __init__(self, config: MyPolicyConfig, dataset_stats: dict[str, Any] = None):
        super().__init__(config, dataset_stats)
        config.validate_features()  # not called automatically by the base class
        self.config = config
        self.model = ...  # your nn.Module here

    def reset(self):
-        """Reset episode state."""
+        """Reset per-episode state. Called by lerobot-eval at the start of each episode."""
        ...

    def get_optim_params(self) -> dict:
@@ -147,35 +129,51 @@ class MyCustomPolicy(PreTrainedPolicy):
        ...

    def select_action(self, batch: dict[str, torch.Tensor], **kwargs) -> torch.Tensor:
-        """Return a single action for the current timestep (called at inference)."""
+        """Return a single action for the current timestep (called every step at inference)."""
        ...

-    def forward(self, batch: dict[str, torch.Tensor]) -> dict[str, torch.Tensor]:
+    def forward(self, batch: dict[str, torch.Tensor]) -> tuple[torch.Tensor, dict | None]:
        """Compute the training loss.

+        Returns `(loss, output_dict)`. `output_dict` may be `None`; everything in it must be
+        logging-friendly Python natives (no tensors with gradients).
+
        `batch["action_is_pad"]` is a bool mask of shape (B, horizon) that marks
-        timesteps padded because the episode ended before `horizon` steps, you
+        timesteps padded because the episode ended before `horizon` steps; you
        can exclude those from your loss.
        """
        actions = batch[ACTION]
        action_is_pad = batch.get("action_is_pad")
        ...
-        return {"loss": ...}
+        return loss, {"some_loss_component": some_loss_component.item()}
 ```

-## Step 4: Add Data Processors
+The methods called by the train/eval loops:

-Create processor functions. For a concrete reference, see [processor_act.py](https://github.com/huggingface/lerobot/blob/main/src/lerobot/policies/act/processor_act.py) or [processor_diffusion.py](https://github.com/huggingface/lerobot/blob/main/src/lerobot/policies/diffusion/processor_diffusion.py).
+| Method                                                            | Used by           | What it does                                                                                                                                                                                                                                         |
+| ----------------------------------------------------------------- | ----------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
+| `reset() -> None`                                                 | `lerobot-eval`    | Clear per-episode state at the start of each episode.                                                                                                                                                                                                |
+| `select_action(batch, **kwargs) -> Tensor`                        | `lerobot-eval`    | Return the next action `(B, action_dim)`. Called every step.                                                                                                                                                                                         |
+| `predict_action_chunk(batch, **kwargs) -> Tensor`                 | the policy itself | Return an action chunk `(B, chunk_size, action_dim)`. Currently abstract on the base class — raise `NotImplementedError` if your policy doesn't chunk.                                                                                               |
+| `forward(batch, reduction="mean") -> tuple[Tensor, dict \| None]` | `lerobot-train`   | Return `(loss, output_dict)`. Accept `reduction="none"` if you want to support per-sample weighting.                                                                                                                                                 |
+| `get_optim_params() -> dict`                                      | the optimizer     | Return `self.parameters()` for simple policies; return a named parameter dict for [multi-optimizer policies](https://github.com/huggingface/lerobot/blob/ecd38c50d7d15b4184cf42649ff1185ee2e11eeb/src/lerobot/policies/sac/modeling_sac.py#L61-L73). |
+| `update() -> None` _(optional)_                                   | `lerobot-train`   | Called after each optimizer step _if defined_. Use for EMA, target nets, replay buffers (TDMPC uses this).                                                                                                                                           |
+
+Batches are flat dictionaries keyed by the constants in [`lerobot.utils.constants`](https://github.com/huggingface/lerobot/blob/main/src/lerobot/utils/constants.py): `OBS_STATE` (`observation.state.<motor>`), `OBS_IMAGES` (`observation.images.<camera>`), `OBS_LANGUAGE`, `ACTION`, etc. Reuse the constants — don't invent new prefixes.
+
+### Processor functions
+
+LeRobot uses `PolicyProcessorPipeline`s to normalize inputs and de-normalize outputs around your policy. For a concrete reference, see [`processor_act.py`](https://github.com/huggingface/lerobot/blob/main/src/lerobot/policies/act/processor_act.py) or [`processor_diffusion.py`](https://github.com/huggingface/lerobot/blob/main/src/lerobot/policies/diffusion/processor_diffusion.py).

 ```python
-# processor_my_custom_policy.py
+# processor_my_policy.py
 from typing import Any
 import torch

 from lerobot.processor import PolicyAction, PolicyProcessorPipeline


-def make_my_custom_policy_pre_post_processors(
+def make_my_policy_pre_post_processors(
    config,
    dataset_stats: dict[str, dict[str, torch.Tensor]] | None = None,
 ) -> tuple[
@@ -187,11 +185,48 @@ def make_my_custom_policy_pre_post_processors(
    return preprocessor, postprocessor
 ```

-**Important - function naming:** LeRobot discovers your processor by name. The function **must** be called `make_{policy_name}_pre_post_processors` (matching the string you passed to `@PreTrainedConfig.register_subclass`).
+**Important — function naming:** LeRobot discovers your processor by name. The function **must** be called `make_{policy_name}_pre_post_processors` (matching the string you passed to `@PreTrainedConfig.register_subclass`).

-## Step 5: Package Initialization
+---

-Expose your classes in the package's `__init__.py`:
+## Path A: Out-of-tree plugin
+
+The fastest way to ship a policy: package it as a standalone Python distribution and install it alongside LeRobot. No PR required, you own the release cycle, and you can publish to PyPI under your own namespace.
+
+### Package structure
+
+Create a package with the prefix `lerobot_policy_` (IMPORTANT!) followed by your policy name:
+
+```bash
+lerobot_policy_my_policy/
+├── pyproject.toml
+└── src/
+    └── lerobot_policy_my_policy/
+        ├── __init__.py
+        ├── configuration_my_policy.py
+        ├── modeling_my_policy.py
+        └── processor_my_policy.py
+```
+
+### `pyproject.toml`
+
+```toml
+[project]
+name = "lerobot_policy_my_policy"
+version = "0.1.0"
+dependencies = [
+    # your policy-specific dependencies
+]
+requires-python = ">= 3.12"
+
+[build-system]
+build-backend = # your-build-backend
+requires = # your-build-system
+```
+
+### Package `__init__.py`
+
+Expose your classes in the package's `__init__.py` and guard against missing `lerobot`:

 ```python
 # __init__.py
@@ -204,44 +239,148 @@ except ImportError:
        "lerobot is not installed. Please install lerobot to use this policy package."
    )

-from .configuration_my_custom_policy import MyCustomPolicyConfig
-from .modeling_my_custom_policy import MyCustomPolicy
-from .processor_my_custom_policy import make_my_custom_policy_pre_post_processors
+from .configuration_my_policy import MyPolicyConfig
+from .modeling_my_policy import MyPolicy
+from .processor_my_policy import make_my_policy_pre_post_processors

 __all__ = [
-    "MyCustomPolicyConfig",
-    "MyCustomPolicy",
-    "make_my_custom_policy_pre_post_processors",
+    "MyPolicyConfig",
+    "MyPolicy",
+    "make_my_policy_pre_post_processors",
 ]
 ```

-## Step 6: Installation and Usage
-
-### Install Your Policy Package
+### Install and use

 ```bash
-cd lerobot_policy_my_custom_policy
+cd lerobot_policy_my_policy
 pip install -e .

 # Or install from PyPI if published
-pip install lerobot_policy_my_custom_policy
+pip install lerobot_policy_my_policy
 ```

-### Use Your Policy
-
 Once installed, your policy automatically integrates with LeRobot's training and evaluation tools:

 ```bash
 lerobot-train \
-    --policy.type my_custom_policy \
+    --policy.type my_policy \
    --env.type pusht \
    --steps 200000
 ```

-## Examples and Community Contributions
+---
+
+## Path B: Contributing in-tree
+
+When your policy has stabilized and there's clear value in shipping it with the library, you can land it directly in LeRobot. Read the general [contribution guide](./contributing) and the [PR template](https://github.com/huggingface/lerobot/blob/main/.github/PULL_REQUEST_TEMPLATE.md) first — that's where you'll find the testing/quality expectations every PR has to meet (`pre-commit run -a`, `pytest`, the community-review rule, etc.). What's below is the policy-specific layer on top of that.
+
+### In-tree layout
+
+```
+src/lerobot/policies/my_policy/
+├── __init__.py                    # re-exports config + modeling + processor factory
+├── configuration_my_policy.py     # MyPolicyConfig + @register_subclass
+├── modeling_my_policy.py          # MyPolicy(PreTrainedPolicy)
+├── processor_my_policy.py         # make_my_policy_pre_post_processors
+└── README.md                      # symlink → ../../../../docs/source/policy_my_policy_README.md
+```
+
+Two notes:
+
+- The `README.md` next to the source is a **symlink** into `docs/source/policy_<name>_README.md` — the actual file lives under `docs/`. Existing policies (act, smolvla, diffusion, …) all do this; copy one of those symlinks. The policy README is conventionally minimal: paper link + BibTeX citation.
+- The user-facing tutorial — what to install, how to train, hyperparameters, benchmark numbers — lives separately at `docs/source/<my_policy>.mdx` and is registered in `_toctree.yml` under "Policies".
+
+The file names are load-bearing: the factory does lazy imports by name, and the processor is discovered by the `make_<policy_name>_pre_post_processors` convention.
+
+### Wiring
+
+Three places need to know about your policy. All by name.
+
+1. **`policies/__init__.py`** — re-export `MyPolicyConfig` and add it to `__all__`. **Don't** re-export the modeling class; it loads lazily through the factory (so `import lerobot` stays fast).
+2. **`factory.py:get_policy_class`** — add a branch returning `MyPolicy` from a lazy import.
+3. **`factory.py:make_policy_config`** and **`factory.py:make_pre_post_processors`** — same idea, two more branches.
+
+Mirror an existing policy that's structurally similar to yours; the diff is small.
+
+### Heavy / optional dependencies
+
+Most policies need a heavy backbone (transformers, diffusers, a specific VLM SDK). The convention is **two-step gating**: a `TYPE_CHECKING`-guarded import at module top, and a `require_package` runtime check in the constructor. [`modeling_diffusion.py`](https://github.com/huggingface/lerobot/blob/main/src/lerobot/policies/diffusion/modeling_diffusion.py) is the canonical reference:
+
+```python
+from typing import TYPE_CHECKING
+from lerobot.utils.import_utils import _diffusers_available, require_package
+
+if TYPE_CHECKING or _diffusers_available:
+    from diffusers.schedulers.scheduling_ddim import DDIMScheduler
+else:
+    DDIMScheduler = None  # keeps the symbol bindable at import time
+
+class DiffusionPolicy(PreTrainedPolicy):
+    def __init__(self, config):
+        require_package("diffusers", extra="diffusion")
+        super().__init__(config)
+        ...
+```
+
+This way:
+
+- `import lerobot.policies` keeps working without the extra installed (the symbol is just bound to `None`).
+- Type checkers see the real symbol.
+- Instantiating the policy without the extra raises a clear `ImportError` pointing at `pip install 'lerobot[diffusion]'`.
+
+Add a matching extra to [`pyproject.toml`](https://github.com/huggingface/lerobot/blob/main/pyproject.toml) `[project.optional-dependencies]` and include it in the `all` extra so `pip install 'lerobot[all]'` keeps installing everything.
+
+### Benchmarks and a published checkpoint
+
+A new policy is much easier to review — and far more useful — when it ships with a working checkpoint and at least one number you can reproduce.
+
+**Pick at least one in-tree benchmark.** LeRobot ships sim benchmarks with per-benchmark Docker images (LIBERO, LIBERO-plus, Meta-World, RoboTwin 2.0, RoboCasa365, RoboCerebra, RoboMME, VLABench and more). Pick the one that matches your policy's modality — VLAs usually go to LIBERO or VLABench; image-only BC to LIBERO or Meta-World. The full list lives under [Benchmarks](./libero) in the docs sidebar.
+
+**Push the checkpoint & processors** to the Hub under `lerobot/<policy>_<benchmark>` (or your namespace if you don't have write access; a maintainer can mirror it). Use `PreTrainedPolicy.push_model_to_hub` so the repo gets `config.json`, `model.safetensors`, and a model card.
+
+**Report results in your policy's MDX**, with the exact `lerobot-eval` command and hardware so anyone can re-run:
+
+```markdown
+## Results
+
+Evaluated on LIBERO with `lerobot/<policy>_libero`:
+
+| Suite          | Success rate | n_episodes |
+| -------------- | -----------: | ---------: |
+| libero_spatial |        87.5% |         50 |
+| libero_object  |        93.0% |         50 |
+| libero_goal    |        81.5% |         50 |
+| libero_10      |        62.0% |         50 |
+| **average**    |    **81.0%** |        200 |
+
+Reproduce: `lerobot-eval --policy.path=lerobot/<policy>_libero --env.type=libero --env.task=libero_spatial --eval.n_episodes=50` (1× A100 40 GB).
+```
+
+Use `n_episodes ≥ 50` per suite for stable success-rate estimates.
+
+If your policy is real-robot-only and no sim benchmark applies, swap the sim eval for: a public training dataset on the Hub, the `lerobot-train` command, the checkpoint, and a real-robot success rate over ≥10 episodes via `lerobot-rollout --policy.path=...`.
+
+### PR checklist
+
+The general expectations are in [`CONTRIBUTING.md`](https://github.com/huggingface/lerobot/blob/main/CONTRIBUTING.md) and the [PR template](https://github.com/huggingface/lerobot/blob/main/.github/PULL_REQUEST_TEMPLATE.md). On top of those, reviewers will look for:
+
+- [ ] `MyPolicy` and `MyPolicyConfig` cover the surface above; `__init_subclass__` accepts the class.
+- [ ] `factory.py` and `policies/__init__.py` are wired (lazy imports for modeling).
+- [ ] `make_my_policy_pre_post_processors` follows the naming convention.
+- [ ] Optional deps live behind a `[project.optional-dependencies]` extra and the `TYPE_CHECKING + require_package` guard.
+- [ ] `tests/policies/` updated; backward-compat artifact committed & policy-specific tests.
+- [ ] `src/lerobot/policies/<name>/README.md` symlinked into `docs/source/policy_<name>_README.md`; user-facing `docs/source/<name>.mdx` written and added to `_toctree.yml`.
+- [ ] At least one reproducible benchmark eval in the policy MDX with a published checkpoint (sim benchmark, or real-robot dataset + checkpoint).
+
+The fastest way to get a clean PR is to copy the directory of the existing policy closest to yours, rename, and replace contents method by method. Don't wait until everything is polished — open a draft PR early and iterate with us; reviewers would much rather give feedback on a half-finished branch than a fully-merged one.
+
+---
+
+## Examples and community contributions

 Check out these example policy implementations:

- [DiTFlow Policy](https://github.com/danielsanjosepro/lerobot_policy_ditflow) - Diffusion Transformer policy with flow-matching objective. Try it out in this example: [DiTFlow Example](https://github.com/danielsanjosepro/test_lerobot_policy_ditflow)
+- [DiTFlow Policy](https://github.com/danielsanjosepro/lerobot_policy_ditflow) — Diffusion Transformer policy with flow-matching objective. Try it out in this example: [DiTFlow Example](https://github.com/danielsanjosepro/test_lerobot_policy_ditflow)

-Share your policy implementations with the community! 🤗
+Thanks for taking the time to bring a new policy into LeRobot. Every architecture that lands in `main` — and every plugin published by the community — makes the library a little more useful for the next person, and a little more representative of where robot learning is going. We're looking forward to seeing what you ship. 🤗
@@ -0,0 +1,98 @@
+# Compute HW Guide for LeRobot Training
+
+Rough sizing for training a LeRobot policy: how much VRAM each policy needs, what training time looks like, and where to run when local hardware isn't enough.
+
+The numbers below are **indicative** — order-of-magnitude figures for picking hardware, not exact predictions. Throughput depends heavily on dataset I/O, image resolution, batch size, and number of GPUs.
+
+## Memory by policy group
+
+Policies cluster by backbone size; the groupings below give a single VRAM envelope per group instead of repeating numbers per policy. Memory scales roughly linearly with batch size; AdamW (the LeRobot default) carries optimizer state that adds ~30–100% over a forward+backward pass alone.
+
+| Group      | Policies                                    | Peak VRAM (BS 8, AdamW) | Suitable starter GPUs             |
+| ---------- | ------------------------------------------- | ----------------------: | --------------------------------- |
+| Light BC   | `act`, `vqbet`, `tdmpc`                     |                  ~2–6GB | Laptop GPU (RTX 3060), L4, A10G   |
+| Diffusion  | `diffusion`, `multi_task_dit`               |                 ~8–14GB | RTX 4070+ / L4 / A10G             |
+| Small VLA  | `smolvla`                                   |                ~10–16GB | RTX 4080+ / L4 / A10G             |
+| Large VLA  | `pi0`, `pi0_fast`, `pi05`, `xvla`, `wall_x` |                ~24–40GB | A100 40 GB+ (24 GB tight at BS 1) |
+| Multimodal | `groot`, `eo1`                              |                ~24–40GB | A100 40 GB+                       |
+| RL         | `sac`                                       |             config-dep. | See [HIL-SERL guide](./hilserl)   |
+
+Memory-bound? Drop the batch size (~linear), use gradient accumulation to recover effective batch, or for SmolVLA leave `freeze_vision_encoder=True`.
+
+## Training time
+
+Robotics imitation learning typically converges in **5–10 epochs over the dataset**, not hundreds of thousands of raw steps. Once you know your epoch count, wall-clock is essentially:
+
+```text
+total_frames    = sum of frames over all episodes      # 50 ep × 30 fps × 30 s ≈ 45,000
+steps_per_epoch = ceil(total_frames / (num_gpus × batch_size))
+total_steps     = epochs × steps_per_epoch
+wall_clock      ≈ total_steps × per_step_time
+```
+
+Per-step time depends on the policy and the GPU. The numbers in the table below are anchors — pick the row closest to your setup and scale linearly with `total_steps` if you train longer or shorter.
+
+### Common scenarios
+
+Indicative wall-clock for **5 epochs on a ~50-episode dataset (~45k frames at 30 fps × 30 s)**, default optimizer (AdamW), 640×480 images:
+
+| Setup                                | Policy         | Batch | Wall-clock |
+| ------------------------------------ | -------------- | ----- | ---------: |
+| Single RTX 4090 / RTX 3090 (24 GB)   | `act`          | 8     |  ~30–60min |
+| Single RTX 4090 / RTX 3090 (24 GB)   | `diffusion`    | 8     |      ~2–4h |
+| Single L4 / A10G (24 GB)             | `act`          | 8     |      ~1–2h |
+| Single L4 / A10G (24 GB)             | `smolvla`      | 4     |      ~3–6h |
+| Single A100 40 GB                    | `smolvla`      | 16    |      ~1–2h |
+| Single A100 40 GB                    | `pi0` / `pi05` | 4     |      ~4–8h |
+| 4× H100 80 GB cluster (`accelerate`) | `diffusion`    | 32    |  ~30–60min |
+| 4× H100 80 GB cluster (`accelerate`) | `smolvla`      | 32    |      ~1–2h |
+| Apple Silicon M1/M2/M3 Max (MPS)     | `act`          | 4     |     ~6–14h |
+
+These are order-of-magnitude figures. Real runs deviate by ±50% depending on image resolution, dataset I/O, dataloader threading, and exact GPU SKU. They are useful as "is this run going to take an hour or a day?" intuition, not as SLAs.
+
+### Multi-GPU matters a lot
+
+`accelerate launch --num_processes=N` is the easiest way to cut training time. Each optimizer step processes `N × batch_size` samples in roughly the same wall-clock as a single-GPU step, so 4 GPUs ≈ 4× speedup for compute-bound runs. See the [Multi GPU training](./multi_gpu_training) guide for the full setup.
+
+Reference data points on a 4×H100 80 GB cluster (`accelerate launch --num_processes=4`), 5000 steps, batch 32, AdamW, dataset [`imstevenpmwork/super_poulain_draft`](https://huggingface.co/datasets/imstevenpmwork/super_poulain_draft) (~50 episodes, ~640×480 images):
+
+| Policy      | Wall-clock | `update_s` | `dataloading_s` | GPU util | Notable flags                                                                                                                  |
+| ----------- | ---------- | ---------: | --------------: | -------- | ------------------------------------------------------------------------------------------------------------------------------ |
+| `diffusion` | 16m 17s    |      0.167 |           0.015 | ~90%     | defaults (training from scratch)                                                                                               |
+| `smolvla`   | 27m 49s    |      0.312 |           0.011 | ~80%     | `--policy.path=lerobot/smolvla_base`, `freeze_vision_encoder=false`, `train_expert_only=false`                                 |
+| `pi05`      | 3h 41m     |      2.548 |           0.014 | ~95%     | `--policy.pretrained_path=lerobot/pi05_base`, `gradient_checkpointing=true`, `dtype=bfloat16`, vision encoder + expert trained |
+
+The `dataloading_s` vs. `update_s` ratio is the diagnostic that matters: when `dataloading_s` approaches `update_s`, more GPUs stop helping — your dataloader is the bottleneck and you should look at `--num_workers`, image resolution, and disk speed before adding compute.
+
+### Schedule and checkpoints
+
+If you shorten training (e.g. 5k–10k steps on a small dataset), also shorten the LR schedule with `--policy.scheduler_decay_steps≈--steps`. Otherwise the LR stays near its peak and never decays. Same for `--save_freq`.
+
+## Where to run
+
+VRAM is the first filter. Within a tier, pick by budget and availability — the `$`–`$$$$` columns are relative; check current pricing on the provider you actually use.
+
+| Class                      | VRAM  | Tier   | Comfortable for                                             |
+| -------------------------- | ----- | ------ | ----------------------------------------------------------- |
+| RTX 3090 / 4090 (consumer) | 24 GB | `$`    | Light BC, Diffusion, SmolVLA. Tight for VLAs at batch 1.    |
+| L4 / A10G (cloud)          | 24 GB | `$–$$` | Same envelope; common on Google Cloud, RunPod, AWS `g5/g6`. |
+| A100 40 GB                 | 40 GB | `$$$`  | Any policy at reasonable batch sizes.                       |
+| A100 80 GB / H100 80 GB    | 80 GB | `$$$$` | Multi-GPU clusters; large batches for VLAs.                 |
+| **CPU only**               | —     | —      | Don't train. Use Colab or rent a GPU.                       |
+
+### Hugging Face Jobs
+
+[Hugging Face Jobs](https://huggingface.co/docs/hub/jobs) lets you run training on managed HF infrastructure, billed by the second. The repo publishes a ready-to-use image: **`huggingface/lerobot-gpu:latest`**, rebuilt **every night at 02:00 UTC from `main`** ([`docker_publish.yml`](https://github.com/huggingface/lerobot/blob/main/.github/workflows/docker_publish.yml)) — so it tracks the current state of the repo, not a tagged release.
+
+```bash
+hf jobs run --flavor a10g-large huggingface/lerobot-gpu:latest \
+  bash -c "nvidia-smi && lerobot-train \
+    --policy.type=act --dataset.repo_id=<USER>/<DATASET> \
+    --policy.repo_id=<USER>/act_<task> --batch_size=8 --steps=50000"
+```
+
+Notes:
+
+- The leading `nvidia-smi` is a quick sanity check that CUDA is visible inside the container — useful to fail fast if the flavor or driver mismatched.
+- The default Job timeout is 30 minutes; pass `--timeout 4h` (or longer) for real training.
+- `--flavor` maps onto the table above: `t4-small`/`t4-medium` (T4, ACT only), `l4x1`/`l4x4` (L4 24 GB), `a10g-small/large/largex2/largex4` (A10G 24 GB scaled out), `a100-large` (A100). For the current full catalogue + pricing see [https://huggingface.co/docs/hub/jobs](https://huggingface.co/docs/hub/jobs).
@@ -0,0 +1,136 @@
+# OMX Follower — Cube Pick And Place Example
+
+This is an example of what is possible to do with LeRobot on a physical setup.
+It is a WIP and being used internally at LeRobot and specific to our setup, but we hope it can be a useful reference for how to use LeRobot APIs and CLIs.
+
+It includes an end-to-end example for the **OMX Follower** robot arm: pick and place a cube dataset, train a policy, and deploy it autonomously.
+
+## Hardware
+
+| Component | Value                                |
+| --------- | ------------------------------------ |
+| Robot     | OMX Follower                         |
+| Cameras   | 2× OpenCV cameras (wrist + top-down) |
+
+## Scripts
+
+| Script                 | Purpose                                                         |
+| ---------------------- | --------------------------------------------------------------- |
+| `reset_environment.py` | Standalone utility: sweep workspace, grab cube, place cube      |
+| `record_grab.py`       | Automated data collection: reset → place → record grab episodes |
+
+## Setup
+
+Make sure you have LeRobot installed in your env. (See [the installation guide](https://huggingface.co/docs/lerobot/installation))
+
+Next, we will declare some environment variables for convenience. Adjust the camera indices and robot port to match your system configuration.
+
+```bash
+export ROBOT_PORT=/dev/ttyACM0
+export TELEOP_PORT=/dev/ttyACM1
+export HF_USERNAME=<your_hf_username>
+export ROBOT_CAMERAS="{ wrist: {type: opencv, index_or_path: 0, width: 640, height: 480, fps: 30, fourcc: MJPG}, top: {type: opencv, index_or_path: 2, width: 640, height: 480, fps: 30, fourcc: MJPG} }"
+```
+
+## Step 1 — Collect Data
+
+```bash
+lerobot-record \
+    --robot.type=omx_follower \
+    --robot.port=$ROBOT_PORT \
+    --robot.id=omx_follower \
+    --robot.cameras="$ROBOT_CAMERAS" \
+    --teleop.type=omx_leader \
+    --teleop.port=$TELEOP_PORT \
+    --teleop.id=omx_leader \
+    --dataset.repo_id=$HF_USERNAME/omx_pickandplace \
+    --dataset.root=data/omx_pickandplace \
+    --dataset.num_episodes=50 \
+    --dataset.single_task="Pick the cube and place it in the blue square" \
+    --dataset.streaming_encoding=true \
+    --dataset.push_to_hub=true
+```
+
+### Bonus Auto-Collect script
+
+/!\ This is specific to our setup and the task of picking and placing a cube. It is not a general-purpose data collection script. As you may notice, it doesn't require a teleop.
+
+```bash
+python -m examples.omx.record_grab \
+    --robot.type=omx_follower \
+    --robot.port=$ROBOT_PORT \
+    --robot.id=omx_follower \
+    --robot.cameras="$ROBOT_CAMERAS" \
+    --dataset.repo_id=$HF_USERNAME/omx_pickandplace \
+    --dataset.root=data/omx_pickandplace \
+    --dataset.num_episodes=50 \
+    --dataset.single_task="Pick the cube and place it in the blue square" \
+    --dataset.streaming_encoding=true \
+    --dataset.push_to_hub=true
+```
+
+Each episode:
+
+1. The arm grabs the cube from the center of the workspace and places it at a random position.
+2. The arm returns to HOME.
+3. A targeted grab is recorded: HOME → approach raised → lower onto cube → grasp → lift → carry → drop → HOME.
+
+A dataset is already available here [`maximellerbach/omx_pickandplace`](https://huggingface.co/datasets/maximellerbach/omx_pickandplace), so you can skip directly to training if you want.
+
+## Step 2 — Train
+
+To train a simple `ACT` policy on the collected dataset, you can use the `lerobot-train` CLI:
+
+```bash
+lerobot-train \
+    --dataset.repo_id=$HF_USERNAME/omx_pickandplace \
+    --policy.type=act \
+    --output_dir=outputs/train/omx_pickandplace_act \
+    --policy.device=cuda \
+    --policy.repo_id=$HF_USERNAME/omx_pickandplace_act \
+    --steps=20000 \
+    --wandb.enable=true
+```
+
+A pretrained `ACT` policy is already available here [`maximellerbach/omx_pickandplace_act`](https://huggingface.co/maximellerbach/omx_pickandplace_act).
+
+## Step 3 — Rollout
+
+Use the `lerobot-rollout` CLI with base strategy:
+
+```bash
+lerobot-rollout \
+    --strategy.type=base \
+    --robot.type=omx_follower \
+    --robot.port=$ROBOT_PORT \
+    --robot.id=omx_follower \
+    --robot.cameras="$ROBOT_CAMERAS" \
+    --policy.path=$HF_USERNAME/omx_pickandplace_act \
+```
+
+For continuous recording with automatic upload (sentry mode):
+
+```bash
+lerobot-rollout \
+    --strategy.type=sentry \
+    --strategy.upload_every_n_episodes=10 \
+    --robot.type=omx_follower \
+    --robot.port=$ROBOT_PORT \
+    --robot.id=omx_follower \
+    --robot.cameras="$ROBOT_CAMERAS" \
+    --policy.path=$HF_USERNAME/omx_pickandplace_act \
+    --dataset.repo_id=$HF_USERNAME/rollout_omx_pickandplace_act \
+```
+
+## Environment Reset Utility
+
+Those are specific to this particular physical setup. Those are scripts that execute hardcoded sequences of actions on the robot to reset the environment, which is useful for data collection and evaluation. They are not general-purpose scripts.
+
+`reset_environment.py` can be run standalone to prepare the workspace:
+
+```bash
+# Grab cube + place it at a random position on the left side
+python -m examples.omx.reset_environment --port $ROBOT_PORT --mode grab_and_place
+```
+
+It also exposes `grab_cube(robot)` and `place_cube(robot)` for use in custom scripts.
@@ -0,0 +1,422 @@
+#!/usr/bin/env python3
+"""
+Auto-record grab episodes for the OMX robot arm.
+
+Each episode cycle:
+  1. grab_and_place  — grab cube from workspace center and place at a random (pan, reach) position
+  2. HOME            — return arm to home with gripper open
+  3. record_grab     — execute a targeted grab to the stored position while recording
+                       observations + actions to a LeRobotDataset
+
+Usage (run from repo root):
+    python -m examples.omx.record_grab \\
+        --robot.type=omx_follower \\
+        --robot.port=/dev/ttyACM0 \\
+        --robot.id=omx_follower \\
+        --robot.cameras="{ wrist: {type: opencv, index_or_path: 6, width: 640, height: 480, fps: 30, fourcc: MJPG}, top: {type: opencv, index_or_path: 4, width: 640, height: 480, fps: 30, fourcc: MJPG} }" \\
+        --dataset.repo_id=<hf_username>/<dataset_name> \\
+        --dataset.root=data/omx_grab \\
+        --dataset.num_episodes=50 \\
+        --dataset.single_task="Grab the cube" \\
+        --dataset.streaming_encoding=true
+"""
+
+import logging
+from dataclasses import dataclass
+from pprint import pformat
+
+import numpy as np
+
+from lerobot.cameras import CameraConfig  # noqa: F401
+from lerobot.cameras.opencv import OpenCVCameraConfig  # noqa: F401
+from lerobot.configs import parser
+from lerobot.configs.dataset import DatasetRecordConfig
+from lerobot.datasets import (
+    LeRobotDataset,
+    VideoEncodingManager,
+    aggregate_pipeline_dataset_features,
+    create_initial_features,
+)
+from lerobot.processor import make_default_processors
+from lerobot.robots import RobotConfig, make_robot_from_config
+from lerobot.robots.omx_follower import OmxFollower
+from lerobot.utils.constants import ACTION, OBS_STR
+from lerobot.utils.feature_utils import build_dataset_frame, combine_feature_dicts
+from lerobot.utils.robot_utils import precise_sleep
+
+from .reset_environment import (
+    APPROACH_SPEED,
+    GRIPPER_CLOSE_POS,
+    HOME_POSE,
+    PUSH_END_ELBOW_FLEX,
+    PUSH_END_SHOULDER_LIFT,
+    PUSH_START_ELBOW_FLEX,
+    PUSH_START_SHOULDER_LIFT,
+    array_to_pose,
+    grab_cube,
+    horizontal_wrist_flex,
+    move_to_pose,
+    place_cube,
+    pose_to_array,
+)
+
+# ── Grab-episode motion parameters ────────────────────────────────────────────
+
+# Shoulder-lift offset for the raised approach phase (subtracted from the target sl, arm is higher).
+GRAB_RAISE_SL_OFFSET = 20.0
+GRAB_LOWER_SPEED = 20.0
+RECORD_SPEED = 30.0
+
+# Pose the arm travels to after closing the gripper (cube held).
+GRAB_CARRY_POSE = {
+    "shoulder_pan.pos": -23.0,
+    "shoulder_lift.pos": 5.0,
+    "elbow_flex.pos": 18.0,
+    "wrist_flex.pos": -14.0,
+    "wrist_roll.pos": 0.0,
+    "gripper.pos": GRIPPER_CLOSE_POS,
+}
+
+# Per-joint jitter limits (degrees) applied to transit waypoints for human-like variation.
+# Cube-approach and carry poses are never jittered to preserve precision.
+_JITTER_LIMITS: dict[str, float] = {
+    "shoulder_pan.pos": 5.0,
+    "shoulder_lift.pos": 4.0,
+    "elbow_flex.pos": 4.0,
+    "wrist_flex.pos": 3.0,
+    "wrist_roll.pos": 2.0,
+    "gripper.pos": 0.0,
+}
+
+
+def _jitter_pose(pose: dict, rng: np.random.Generator) -> dict:
+    """Return a copy of pose with independent per-joint random perturbations."""
+    return {
+        k: v + rng.uniform(-_JITTER_LIMITS.get(k, 0.0), _JITTER_LIMITS.get(k, 0.0)) for k, v in pose.items()
+    }
+
+
+def _random_stuck_pose(rng: np.random.Generator) -> dict:
+    """Return a physically plausible stuck pose (failed grasp), gripper closed.
+
+    ef bounds are piecewise-linear in sl so the arm stays in a reachable,
+    table-safe envelope across the full sl range:
+      sl=-50 → ef ∈ [  0,  50]   (arm raised, can be bent forward)
+      sl=  0 → ef ∈ [-25,  25]   (mid reach)
+      sl= 30 → ef ∈ [-20,   0]   (arm extended, little room to flex)
+    wrist_flex is randomly offset from the horizontal value.
+    """
+    pan = float(rng.uniform(-5.0, 35.0))
+    sl = float(rng.uniform(-50.0, 30.0))
+
+    if sl <= 0.0:
+        alpha = (sl + 50.0) / 50.0  # 0 at sl=-50, 1 at sl=0
+        ef_lo = alpha * -25.0  # 0 → -25
+        ef_hi = 50.0 + alpha * -25.0  # 50 → 25
+    else:
+        alpha = sl / 30.0  # 0 at sl=0, 1 at sl=30
+        ef_lo = -25.0 + alpha * 5.0  # -25 → -20
+        ef_hi = 25.0 + alpha * -25.0  # 25 → 0
+
+    ef = float(rng.uniform(ef_lo, ef_hi))
+    wf = horizontal_wrist_flex(sl, ef) + float(rng.uniform(-15.0, 15.0))
+    return {
+        "shoulder_pan.pos": pan,
+        "shoulder_lift.pos": sl,
+        "elbow_flex.pos": ef,
+        "wrist_flex.pos": wf,
+        "wrist_roll.pos": float(rng.uniform(-15.0, 15.0)),
+        "gripper.pos": GRIPPER_CLOSE_POS,
+    }
+
+
+logger = logging.getLogger(__name__)
+
+
+@dataclass
+class OmxRecordGrabConfig:
+    robot: RobotConfig
+    dataset: DatasetRecordConfig
+    # Resume recording on an existing dataset.
+    resume: bool = False
+    # Fraction of episodes that start from a random stuck pose (gripper closed) to
+    # generate recovery data.  0.0 = disabled, 1.0 = all episodes are recovery starts.
+    recovery_prob: float = 0.5
+
+
+def record_episode_spline(
+    robot: OmxFollower,
+    waypoints: list[dict],
+    speeds: list[float],
+    dataset: LeRobotDataset,
+    task: str,
+) -> None:
+    """Execute a Catmull-Rom-style spline through waypoints, recording each frame.
+
+    Segment durations are parameterized from the maximum absolute joint delta
+    between consecutive waypoints divided by the requested segment speed,
+    producing non-uniform timing in joint space. Interior tangents are derived
+    from the adjacent per-segment velocities, with clamped (zero-velocity)
+    endpoints so the arm starts and stops smoothly. Each segment is cubic
+    Hermite, giving C1 continuity at every waypoint.
+    """
+    pts = [pose_to_array(w) for w in waypoints]
+    n = len(pts)
+
+    # Steps and duration per segment
+    n_steps_list = []
+    timestamps = []
+    for i in range(n - 1):
+        max_dist = float(np.max(np.abs(pts[i + 1] - pts[i])))
+        ns = max(1, int(max_dist / speeds[i] * dataset.fps)) if max_dist >= 0.5 else 0
+        n_steps_list.append(ns)
+        timestamps.append(ns / dataset.fps)
+
+    # Velocity tangents (deg/sec) — clamped at endpoints, Catmull-Rom for interior
+    vels = [np.zeros_like(pts[0])]
+    for i in range(1, n - 1):
+        v_prev = (pts[i] - pts[i - 1]) / timestamps[i - 1] if timestamps[i - 1] > 0 else np.zeros_like(pts[0])
+        v_next = (pts[i + 1] - pts[i]) / timestamps[i] if timestamps[i] > 0 else np.zeros_like(pts[0])
+        vels.append(0.5 * (v_prev + v_next))
+    vels.append(np.zeros_like(pts[0]))
+
+    dt = 1.0 / dataset.fps
+    for seg in range(n - 1):
+        ns = n_steps_list[seg]
+        if ns == 0:
+            continue
+        p0, p1 = pts[seg], pts[seg + 1]
+        # Scale velocity (deg/sec) to t-space tangent (deg/t-unit, where t: 0→1 over ns steps)
+        m0 = vels[seg] * timestamps[seg]
+        m1 = vels[seg + 1] * timestamps[seg]
+
+        for step in range(1, ns + 1):
+            t = step / ns
+            h00 = 2 * t**3 - 3 * t**2 + 1
+            h10 = t**3 - 2 * t**2 + t
+            h01 = -2 * t**3 + 3 * t**2
+            h11 = t**3 - t**2
+            commanded = h00 * p0 + h10 * m0 + h01 * p1 + h11 * m1
+
+            action = array_to_pose(commanded)
+            robot.send_action(action)
+            obs = robot.get_observation()
+            obs_frame = build_dataset_frame(dataset.features, obs, prefix=OBS_STR)
+            action_frame = build_dataset_frame(dataset.features, action, prefix=ACTION)
+            dataset.add_frame({**obs_frame, **action_frame, "task": task})
+            precise_sleep(dt)
+
+
+def record_grab_episode(
+    robot: OmxFollower,
+    dataset: LeRobotDataset,
+    pan: float,
+    t: float,
+    task: str,
+    recovery_start: bool = False,
+) -> None:
+    """Execute a targeted grab to the stored (pan, t) position, recording every frame.
+
+    Normal sequence (initial HOME move is NOT recorded):
+      HOME → raised approach above cube → lower → close gripper
+           → raise [jittered] → retract [jittered] → GRAB_CARRY_POSE → drop → HOME
+
+    Recovery sequence (recovery_start=True): arm is moved to a random stuck pose
+    (gripper closed) without recording, then recording begins from there:
+      stuck_pose → raised approach above cube → [normal grab sequence from there]
+
+    All segments are joined by a Catmull-Rom spline (C1-continuous velocities).
+    """
+    sl = PUSH_START_SHOULDER_LIFT + t * (PUSH_END_SHOULDER_LIFT - PUSH_START_SHOULDER_LIFT)
+    ef = PUSH_START_ELBOW_FLEX + t * (PUSH_END_ELBOW_FLEX - PUSH_START_ELBOW_FLEX)
+    sl_raised = sl - GRAB_RAISE_SL_OFFSET
+    wf_horizontal = horizontal_wrist_flex(sl, ef)
+
+    rng = np.random.default_rng()
+
+    if recovery_start:
+        stuck_pose = _random_stuck_pose(rng)
+        logger.info(f"Recovery start: {stuck_pose}")
+        move_to_pose(robot, stuck_pose, APPROACH_SPEED)
+        first_waypoints = [stuck_pose]
+        first_speeds = []
+    else:
+        jittery_start = _jitter_pose(HOME_POSE, rng)
+        move_to_pose(robot, jittery_start, APPROACH_SPEED)
+        first_waypoints = [jittery_start]
+        first_speeds = []
+
+    waypoints = first_waypoints + [
+        {  # raised approach: arm above cube
+            "shoulder_pan.pos": pan,
+            "shoulder_lift.pos": sl_raised,
+            "elbow_flex.pos": ef,
+            "wrist_flex.pos": horizontal_wrist_flex(sl_raised, ef),
+            "wrist_roll.pos": 0.0,
+            "gripper.pos": 60.0,
+        },
+        {  # lower onto cube — no jitter: precision needed
+            "shoulder_pan.pos": pan,
+            "shoulder_lift.pos": sl,
+            "elbow_flex.pos": ef,
+            "wrist_flex.pos": wf_horizontal,
+            "wrist_roll.pos": 0.0,
+            "gripper.pos": 60.0,
+        },
+        {  # close gripper — no jitter: precision needed
+            "shoulder_pan.pos": pan,
+            "shoulder_lift.pos": sl,
+            "elbow_flex.pos": ef,
+            "wrist_flex.pos": wf_horizontal,
+            "wrist_roll.pos": 0.0,
+            "gripper.pos": GRIPPER_CLOSE_POS,
+        },
+        _jitter_pose(
+            {  # raise with cube
+                "shoulder_pan.pos": pan,
+                "shoulder_lift.pos": sl_raised,
+                "elbow_flex.pos": ef,
+                "wrist_flex.pos": horizontal_wrist_flex(sl_raised, ef),
+                "wrist_roll.pos": 0.0,
+                "gripper.pos": GRIPPER_CLOSE_POS,
+            },
+            rng,
+        ),
+        _jitter_pose(
+            {  # retract: fold arm toward HOME before sweeping to carry zone
+                "shoulder_pan.pos": pan * 0.25,
+                "shoulder_lift.pos": HOME_POSE["shoulder_lift.pos"] + 5.0,
+                "elbow_flex.pos": HOME_POSE["elbow_flex.pos"] - 5.0,
+                "wrist_flex.pos": 0.0,
+                "wrist_roll.pos": 0.0,
+                "gripper.pos": GRIPPER_CLOSE_POS,
+            },
+            rng,
+        ),
+        GRAB_CARRY_POSE,  # no jitter: target drop zone
+        {**GRAB_CARRY_POSE, "gripper.pos": 60.0},  # drop cube
+        HOME_POSE,
+    ]
+    speeds = first_speeds + [
+        RECORD_SPEED,  # (HOME →) raised approach
+        GRAB_LOWER_SPEED,  # raised approach → lower
+        GRAB_LOWER_SPEED,  # lower → close gripper
+        RECORD_SPEED,  # close gripper → raise
+        RECORD_SPEED,  # raise → retract
+        RECORD_SPEED,  # retract → carry pose
+        RECORD_SPEED,  # carry pose → drop
+        RECORD_SPEED,  # drop → HOME
+    ]
+
+    record_episode_spline(robot, waypoints, speeds, dataset, task)
+
+    # Dwell at HOME for ~0.5 s before next episode
+    home_action = build_dataset_frame(dataset.features, HOME_POSE, prefix=ACTION)
+    dt = 1.0 / dataset.fps
+    for _ in range(int(dataset.fps * 0.5)):
+        robot.send_action(HOME_POSE)
+        obs = robot.get_observation()
+        obs_frame = build_dataset_frame(dataset.features, obs, prefix=OBS_STR)
+        dataset.add_frame({**obs_frame, **home_action, "task": task})
+        precise_sleep(dt)
+
+
+@parser.wrap()
+def record_grab(cfg: OmxRecordGrabConfig) -> LeRobotDataset:
+    logging.basicConfig(level=logging.INFO, format="%(levelname)s: %(message)s")
+    logger.info(pformat(cfg))
+
+    robot = make_robot_from_config(cfg.robot)
+    use_videos = cfg.dataset.video
+
+    teleop_action_processor, _, robot_obs_processor = make_default_processors()
+
+    dataset_features = combine_feature_dicts(
+        aggregate_pipeline_dataset_features(
+            pipeline=teleop_action_processor,
+            initial_features=create_initial_features(action=robot.action_features),
+            use_videos=use_videos,
+        ),
+        aggregate_pipeline_dataset_features(
+            pipeline=robot_obs_processor,
+            initial_features=create_initial_features(observation=robot.observation_features),
+            use_videos=use_videos,
+        ),
+    )
+
+    num_cameras = len(robot.cameras) if hasattr(robot, "cameras") else 0
+    dataset = None
+
+    try:
+        if cfg.resume:
+            dataset = LeRobotDataset.resume(
+                cfg.dataset.repo_id,
+                root=cfg.dataset.root,
+                streaming_encoding=cfg.dataset.streaming_encoding,
+                batch_encoding_size=cfg.dataset.video_encoding_batch_size,
+                vcodec=cfg.dataset.vcodec,
+                encoder_threads=cfg.dataset.encoder_threads,
+                image_writer_processes=cfg.dataset.num_image_writer_processes if num_cameras > 0 else 0,
+                image_writer_threads=cfg.dataset.num_image_writer_threads_per_camera * num_cameras
+                if num_cameras > 0
+                else 0,
+            )
+        else:
+            cfg.dataset.stamp_repo_id()
+            dataset = LeRobotDataset.create(
+                cfg.dataset.repo_id,
+                cfg.dataset.fps,
+                root=cfg.dataset.root,
+                robot_type=robot.name,
+                features=dataset_features,
+                use_videos=use_videos,
+                streaming_encoding=cfg.dataset.streaming_encoding,
+                batch_encoding_size=cfg.dataset.video_encoding_batch_size,
+                vcodec=cfg.dataset.vcodec,
+                encoder_threads=cfg.dataset.encoder_threads,
+                image_writer_processes=cfg.dataset.num_image_writer_processes if num_cameras > 0 else 0,
+                image_writer_threads=cfg.dataset.num_image_writer_threads_per_camera * num_cameras
+                if num_cameras > 0
+                else 0,
+            )
+
+        robot.connect(calibrate=True)
+
+        rng = np.random.default_rng()
+        with VideoEncodingManager(dataset):
+            for episode_idx in range(cfg.dataset.num_episodes):
+                logger.info(f"=== Episode {episode_idx + 1}/{cfg.dataset.num_episodes} ===")
+
+                logger.info("Step 1: grabbing and placing cube...")
+                grab_cube(robot)
+                pan, t = place_cube(robot)
+                logger.info(f"Cube placed at pan={pan:.1f}, reach={t:.2f}")
+
+                recovery_start = cfg.recovery_prob > 0 and float(rng.random()) < cfg.recovery_prob
+                logger.info(f"Step 2: recording {'recovery ' if recovery_start else ''}grab episode...")
+                record_grab_episode(
+                    robot,
+                    dataset,
+                    pan,
+                    t,
+                    cfg.dataset.single_task,
+                    recovery_start=recovery_start,
+                )
+
+                dataset.save_episode()
+                logger.info(f"Episode {episode_idx + 1} saved.")
+
+    finally:
+        if dataset:
+            dataset.finalize()
+        if robot.is_connected:
+            robot.disconnect()
+
+    if cfg.dataset.push_to_hub and dataset and dataset.num_episodes > 0:
+        dataset.push_to_hub(tags=cfg.dataset.tags, private=cfg.dataset.private)
+
+    return dataset
+
+
+if __name__ == "__main__":
+    record_grab()
@@ -0,0 +1,267 @@
+#!/usr/bin/env python3
+"""
+Auto-reset and cube-grab utility for the OMX robot arm.
+
+Provides:
+  - grab_cube(robot): sweep workspace, center cube, close gripper
+  - place_cube(robot): carry cube to a random position, release
+
+Standalone usage (run from repo root):
+    python -m examples.omx.reset_environment --port /dev/ttyACM1 --mode grab
+    python -m examples.omx.reset_environment --port /dev/ttyACM1 --mode grab_and_place
+
+Joint range: -100 to 100 for arm joints; gripper: 50 = closed, 80 = open.
+
+To read current joint values for calibration, add after robot.connect():
+    obs = robot.get_observation()
+    print({k: round(obs[k], 1) for k in JOINT_NAMES})
+    robot.disconnect(); raise SystemExit
+
+Parallel-to-ground IK: wrist_flex = WRIST_HORIZONTAL_OFFSET - shoulder_lift - elbow_flex.
+Linear interpolation preserves this constraint between any two poses that satisfy it.
+"""
+
+import argparse
+import logging
+
+import numpy as np
+
+from lerobot.robots.omx_follower import OmxFollower, OmxFollowerConfig
+from lerobot.robots.robot import Robot
+from lerobot.utils.robot_utils import precise_sleep
+
+logger = logging.getLogger(__name__)
+
+# ── Poses ─────────────────────────────────────────────────────────────────────
+
+HOME_POSE = {
+    "shoulder_pan.pos": 0.0,
+    "shoulder_lift.pos": -50.0,
+    "elbow_flex.pos": 50.0,
+    "wrist_flex.pos": 0.0,
+    "wrist_roll.pos": 0.0,
+    "gripper.pos": 60.0,
+}
+
+SWEEP_WAYPOINTS = [
+    {
+        "shoulder_pan.pos": -60.0,
+        "shoulder_lift.pos": 50.0,
+        "elbow_flex.pos": -60.0,
+        "wrist_flex.pos": -20.0,
+        "wrist_roll.pos": 0.0,
+        "gripper.pos": 60.0,
+    },
+    {
+        "shoulder_pan.pos": -30.0,
+        "shoulder_lift.pos": 50.0,
+        "elbow_flex.pos": -60.0,
+        "wrist_flex.pos": -5.0,
+        "wrist_roll.pos": 0.0,
+        "gripper.pos": 60.0,
+    },
+    {
+        "shoulder_pan.pos": 20.0,
+        "shoulder_lift.pos": 50.0,
+        "elbow_flex.pos": -55.0,
+        "wrist_flex.pos": -5.0,
+        "wrist_roll.pos": 0.0,
+        "gripper.pos": 60.0,
+    },
+]
+
+# ── Motion parameters ─────────────────────────────────────────────────────────
+
+CONTROL_HZ = 30
+APPROACH_SPEED = 50.0
+SWEEP_SPEED = 40.0
+
+# ── Grab-sequence parameters ──────────────────────────────────────────────────
+
+GRAB_PAN = 0.0
+SWEEP_LEFT_PAN = -60.0
+SWEEP_RIGHT_PAN = 60.0
+SWEEP_END_OFFSET = 5.0  # stop before center so the cube isn't pushed past GRAB_PAN
+SWEEP_END_PAN_RANGE = (15.0, 20.0)
+
+SWEEP_LOW_SHOULDER_LIFT = 50.0
+SWEEP_LOW_ELBOW_FLEX_START = -60.0
+SWEEP_LOW_ELBOW_FLEX_END = -55.0
+
+SWEEP_HIGH_WRIST_FLEX = -20.0  # wrist tilted up during high approach to clear obstacles
+
+PUSH_START_SHOULDER_LIFT = 0.0
+PUSH_START_ELBOW_FLEX = 45.0
+PUSH_END_SHOULDER_LIFT = 50.0
+PUSH_END_ELBOW_FLEX = -50.0
+# Subtracted from shoulder_lift during the push sweep to clear the platform surface.
+# Does not affect the grab-target interpolation in record_grab.py.
+PUSH_RAISE_OFFSET = 5.0
+
+WRIST_HORIZONTAL_OFFSET = 0.0  # tune if gripper tilts during push: + tilts nose up, - down
+GRIPPER_CLOSE_POS = 50.0
+
+PLACE_LEFT_PAN_RANGE = (5.0, 30.0)  # random pan range for cube placement on the left side
+PLACE_REACH_RANGE = (0.1, 0.7)  # 0 = arm retracted (PUSH_START), 1 = fully extended (PUSH_END)
+
+JOINT_NAMES = [
+    "shoulder_pan.pos",
+    "shoulder_lift.pos",
+    "elbow_flex.pos",
+    "wrist_flex.pos",
+    "wrist_roll.pos",
+    "gripper.pos",
+]
+
+# ── Helpers ───────────────────────────────────────────────────────────────────
+
+
+def pose_to_array(pose: dict) -> np.ndarray:
+    return np.array([pose[k] for k in JOINT_NAMES])
+
+
+def array_to_pose(arr: np.ndarray) -> dict:
+    return {k: float(arr[i]) for i, k in enumerate(JOINT_NAMES)}
+
+
+def horizontal_wrist_flex(shoulder_lift: float, elbow_flex: float) -> float:
+    return WRIST_HORIZONTAL_OFFSET - shoulder_lift - elbow_flex
+
+
+def _low_sweep_pose(pan: float, elbow_flex: float, wrist_flex: float | None = None) -> dict:
+    sl = SWEEP_LOW_SHOULDER_LIFT
+    return {
+        "shoulder_pan.pos": pan,
+        "shoulder_lift.pos": sl,
+        "elbow_flex.pos": elbow_flex,
+        "wrist_flex.pos": horizontal_wrist_flex(sl, elbow_flex) if wrist_flex is None else wrist_flex,
+        "wrist_roll.pos": 0.0,
+        "gripper.pos": 60.0,
+    }
+
+
+def _high_sweep_pose(pan: float) -> dict:
+    return {**HOME_POSE, "shoulder_pan.pos": pan, "wrist_flex.pos": SWEEP_HIGH_WRIST_FLEX}
+
+
+def _push_pose(shoulder_lift: float, elbow_flex: float, pan: float = GRAB_PAN, gripper: float = 70.0) -> dict:
+    return {
+        "shoulder_pan.pos": pan,
+        "shoulder_lift.pos": shoulder_lift,
+        "elbow_flex.pos": elbow_flex,
+        "wrist_flex.pos": horizontal_wrist_flex(shoulder_lift, elbow_flex),
+        "wrist_roll.pos": 0.0,
+        "gripper.pos": gripper,
+    }
+
+
+def move_to_pose(robot: Robot, target: dict, speed: float) -> None:
+    """Interpolate from current position to target at the given speed (units/s)."""
+    obs = robot.get_observation()
+    current = np.array([obs[k] for k in JOINT_NAMES])
+    goal = pose_to_array(target)
+
+    max_distance = float(np.max(np.abs(goal - current)))
+    if max_distance < 0.5:
+        return
+
+    n_steps = max(1, int(max_distance / speed * CONTROL_HZ))
+    dt = 1.0 / CONTROL_HZ
+    for step in range(1, n_steps + 1):
+        t = step / n_steps
+        robot.send_action(array_to_pose(current + t * (goal - current)))
+        precise_sleep(dt)
+
+
+# ── Sequences ─────────────────────────────────────────────────────────────────
+
+
+def grab_cube(robot: Robot) -> None:
+    """Left sweep → right sweep → extend arm parallel to ground → close gripper."""
+    move_to_pose(robot, HOME_POSE, APPROACH_SPEED)
+
+    for pan, end_pan in [
+        (SWEEP_LEFT_PAN, GRAB_PAN - SWEEP_END_OFFSET),
+        (SWEEP_RIGHT_PAN, GRAB_PAN + SWEEP_END_OFFSET),
+    ]:
+        logger.info(f"Sweeping {'left' if pan < 0 else 'right'} → center...")
+        move_to_pose(robot, _high_sweep_pose(pan), APPROACH_SPEED)
+        move_to_pose(
+            robot, _low_sweep_pose(pan, SWEEP_LOW_ELBOW_FLEX_START, wrist_flex=-20.0), APPROACH_SPEED
+        )
+        move_to_pose(robot, _low_sweep_pose(end_pan, SWEEP_LOW_ELBOW_FLEX_END, wrist_flex=0.0), SWEEP_SPEED)
+        move_to_pose(robot, HOME_POSE, APPROACH_SPEED)
+
+    logger.info("Extending to push cube into gripper...")
+    move_to_pose(
+        robot,
+        _push_pose(PUSH_START_SHOULDER_LIFT - PUSH_RAISE_OFFSET, PUSH_START_ELBOW_FLEX),
+        APPROACH_SPEED,
+    )
+    move_to_pose(
+        robot,
+        _push_pose(PUSH_END_SHOULDER_LIFT - PUSH_RAISE_OFFSET, PUSH_END_ELBOW_FLEX),
+        SWEEP_SPEED,
+    )
+
+    logger.info("Closing gripper...")
+    move_to_pose(
+        robot,
+        _push_pose(PUSH_END_SHOULDER_LIFT, PUSH_END_ELBOW_FLEX, gripper=GRIPPER_CLOSE_POS),
+        APPROACH_SPEED,
+    )
+
+    logger.info("Grab complete.")
+
+
+def place_cube(robot: Robot) -> tuple[float, float]:
+    """Carry the cube (gripper closed) to a random position on the left side, then release.
+
+    Returns:
+        (pan, t): pan angle and reach scalar [0, 1] of the placement position.
+    """
+    pan = float(np.random.uniform(*PLACE_LEFT_PAN_RANGE))
+    t = float(np.random.uniform(*PLACE_REACH_RANGE))
+    sl = PUSH_START_SHOULDER_LIFT + t * (PUSH_END_SHOULDER_LIFT - PUSH_START_SHOULDER_LIFT)
+    ef = PUSH_START_ELBOW_FLEX + t * (PUSH_END_ELBOW_FLEX - PUSH_START_ELBOW_FLEX)
+    logger.info(f"Placing cube at pan={pan:.1f}, reach={t:.2f}...")
+
+    move_to_pose(robot, {**HOME_POSE, "gripper.pos": GRIPPER_CLOSE_POS}, APPROACH_SPEED)
+    move_to_pose(
+        robot, {**HOME_POSE, "shoulder_pan.pos": pan, "gripper.pos": GRIPPER_CLOSE_POS}, APPROACH_SPEED
+    )
+    move_to_pose(robot, _push_pose(sl, ef, pan=pan, gripper=GRIPPER_CLOSE_POS), APPROACH_SPEED)
+    move_to_pose(robot, _push_pose(sl, ef, pan=pan, gripper=80.0), APPROACH_SPEED)
+    move_to_pose(robot, HOME_POSE, APPROACH_SPEED)
+    logger.info("Place complete.")
+    return pan, t
+
+
+# ── Entry point ───────────────────────────────────────────────────────────────
+
+
+def main():
+    parser = argparse.ArgumentParser(description="OMX arm reset / grab script")
+    parser.add_argument("--port", default="/dev/ttyACM1")
+    parser.add_argument("--robot_id", default="omx_follower")
+    parser.add_argument("--mode", choices=["grab", "grab_and_place"], default="grab_and_place")
+    args = parser.parse_args()
+
+    logging.basicConfig(level=logging.INFO, format="%(levelname)s: %(message)s")
+
+    robot = OmxFollower(OmxFollowerConfig(port=args.port, id=args.robot_id))
+    robot.connect(calibrate=True)
+
+    try:
+        if args.mode == "grab":
+            grab_cube(robot)
+        elif args.mode == "grab_and_place":
+            grab_cube(robot)
+            place_cube(robot)
+
+    finally:
+        robot.disconnect()
+
+
+if __name__ == "__main__":
+    main()
@@ -59,8 +59,8 @@ keywords = ["lerobot", "huggingface", "robotics",  "machine learning", "artifici

 dependencies = [
    # Core ML
-    "torch>=2.7,<2.11.0",
-    "torchvision>=0.22.0,<0.26.0",
+    "torch>=2.7,<2.13.0",
+    "torchvision>=0.22.0,<0.28.0",
    "numpy>=2.0.0,<2.3.0", # NOTE: Explicitly listing numpy helps the resolver converge faster. Upper bound imposed by opencv-python-headless.
    "opencv-python-headless>=4.9.0,<4.14.0",
    "Pillow>=10.0.0,<13.0.0",
@@ -99,7 +99,7 @@ dataset = [
    "pandas>=2.0.0,<3.0.0", # NOTE: Transitive dependency of datasets
    "pyarrow>=21.0.0,<30.0.0", # NOTE: Transitive dependency of datasets
    "lerobot[av-dep]",
-    "torchcodec>=0.3.0,<0.11.0; sys_platform != 'win32' and (sys_platform != 'linux' or (platform_machine != 'aarch64' and platform_machine != 'arm64' and platform_machine != 'armv7l')) and (sys_platform != 'darwin' or platform_machine != 'x86_64')", # NOTE: Windows support starts at version 0.7 (needs torch==2.8), ffmpeg>=8 support starts at version 0.8.1 (needs torch==2.9), system-wide ffmpeg support starts at version 0.10 (needs torch==2.10).
+    "torchcodec>=0.3.0,<0.13.0; sys_platform != 'win32' and (sys_platform != 'linux' or (platform_machine != 'aarch64' and platform_machine != 'arm64' and platform_machine != 'armv7l')) and (sys_platform != 'darwin' or platform_machine != 'x86_64')", # NOTE: Windows support starts at version 0.7 (needs torch==2.8), ffmpeg>=8 support starts at version 0.8.1 (needs torch==2.9), system-wide ffmpeg support starts at version 0.10 (needs torch==2.10), 0.11 needs torch==2.11, 0.12 needs torch==2.12.
    "jsonlines>=4.0.0,<5.0.0",
 ]
 training = [
@@ -256,7 +256,9 @@ class TrainPipelineConfig(HubMixin):
                ) from e

        cli_args = kwargs.pop("cli_args", [])
-        if config_file is not None:
+        # Legacy RA-BC migration only applies to framework-saved checkpoints (always JSON).
+        # Hand-written YAML/TOML configs are expected to use the current sample_weighting schema.
+        if config_file is not None and config_file.endswith(".json"):
            with open(config_file) as f:
                config = json.load(f)
            migrated_config = _migrate_legacy_rabc_fields(config)
@@ -282,7 +282,11 @@ class VideoDecoderCache:
        with self._lock:
            if video_path not in self._cache:
                file_handle = fsspec.open(video_path).__enter__()
-                decoder = VideoDecoder(file_handle, seek_mode="approximate")
+                try:
+                    decoder = VideoDecoder(file_handle, seek_mode="approximate")
+                except Exception:
+                    file_handle.close()
+                    raise
                self._cache[video_path] = (decoder, file_handle)

            return self._cache[video_path][0]
@@ -100,8 +100,8 @@ class DiffusionConfig(PreTrainedConfig):

    # Inputs / output structure.
    n_obs_steps: int = 2
-    horizon: int = 16
-    n_action_steps: int = 8
+    horizon: int = 64
+    n_action_steps: int = 32

    normalization_mapping: dict[str, NormalizationMode] = field(
        default_factory=lambda: {
@@ -122,10 +122,10 @@ class DiffusionConfig(PreTrainedConfig):
    crop_ratio: float = 1.0
    crop_shape: tuple[int, int] | None = None
    crop_is_random: bool = True
-    pretrained_backbone_weights: str | None = None
-    use_group_norm: bool = True
+    pretrained_backbone_weights: str | None = "ResNet18_Weights.IMAGENET1K_V1"
+    use_group_norm: bool = False
    spatial_softmax_num_keypoints: int = 32
-    use_separate_rgb_encoder_per_camera: bool = False
+    use_separate_rgb_encoder_per_camera: bool = True
    # Unet.
    down_dims: tuple[int, ...] = (512, 1024, 2048)
    kernel_size: int = 5
@@ -97,8 +97,8 @@ class VQBeTConfig(PreTrainedConfig):
    vision_backbone: str = "resnet18"
    crop_shape: tuple[int, int] | None = (84, 84)
    crop_is_random: bool = True
-    pretrained_backbone_weights: str | None = None
-    use_group_norm: bool = True
+    pretrained_backbone_weights: str | None = "ResNet18_Weights.IMAGENET1K_V1"
+    use_group_norm: bool = False
    spatial_softmax_num_keypoints: int = 32
    # VQ-VAE
    n_vqvae_training_steps: int = 20000
@@ -939,7 +939,7 @@ class Qwen2_5_VLFlashAttention2(Qwen2_5_VLAttention):
        input_dtype = query_states.dtype
        if input_dtype == torch.float32:
            if torch.is_autocast_enabled():
-                target_dtype = torch.get_autocast_gpu_dtype()
+                target_dtype = torch.get_autocast_dtype(query_states.device.type)
            # Handle the case where the model is quantized
            elif hasattr(self.config, "_pre_quantization_dtype"):
                target_dtype = self.config._pre_quantization_dtype
@@ -985,7 +985,7 @@ class Florence2FlashAttention2(Florence2Attention):
        input_dtype = query_states.dtype
        if input_dtype == torch.float32:
            if torch.is_autocast_enabled():
-                target_dtype = torch.get_autocast_gpu_dtype()
+                target_dtype = torch.get_autocast_dtype(query_states.device.type)
            # Handle the case where the model is quantized
            elif hasattr(self.config, "_pre_quantization_dtype"):
                target_dtype = self.config._pre_quantization_dtype
@@ -46,7 +46,7 @@ class LeKiwiConfig(RobotConfig):
    cameras: dict[str, CameraConfig] = field(default_factory=lekiwi_cameras_config)

    # Set to `True` for backward compatibility with previous policies/dataset
-    use_degrees: bool = False
+    use_degrees: bool = True


@dataclass
@@ -1,3 +1,3 @@
 version https://git-lfs.github.com/spec/v1
-oid sha256:54aecbc1af72a4cd5e9261492f5e7601890517516257aacdf2a0ffb3ce281f1b
+oid sha256:51effd76b73e972f10d31f5084ab906386134b600c87b2668767d30232a902bd
 size 992
@@ -1,3 +1,3 @@
 version https://git-lfs.github.com/spec/v1
-oid sha256:88a9c3775a2aa1e90a08850521970070a4fcf0f6b82aab43cd8ccc5cf77e0013
-size 47424
+oid sha256:d4d7a16ca67f9adefac0e0620a7b2e9c822f2db42faaaced7a89fbad60e5ead4
+size 47680
@@ -1,3 +1,3 @@
 version https://git-lfs.github.com/spec/v1
-oid sha256:91a2635e05a75fe187a5081504c5f35ce3417378813fa2deaf9ca4e8200e1819
+oid sha256:796c439ee8a64bf9901ff8325e7419bda8bd316360ee95e6304e8e1ae0f4c36c
 size 68
@@ -1,3 +1,3 @@
 version https://git-lfs.github.com/spec/v1
-oid sha256:645bff922ac7bea63ad018ebf77c303c0e4cd2c1c0dc5ef3192865281bef3dc6
-size 47424
+oid sha256:ad33a8b47c39c2e1374567ff9da43cdb95e2dbe904c1b02a35051346d3043095
+size 47680
Author	SHA1	Message	Date
Maxime Ellerbach	6d269b28c8	docs(omx): adding some examples and scripts (#3566 ) * docs(omx): adding some examples and scripts * cleaning up and reviewing the cli args * adding __init__.py to example folder, adjusting the examples * adding reference to pretrained act policy * moving `.send_action` before `dataset.add_frame` for consistency Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> Signed-off-by: Maxime Ellerbach <maxime@ellerbach.net> * adjusting docstring Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> Signed-off-by: Maxime Ellerbach <maxime@ellerbach.net> * adressing hardcoded dataset fps * removed init as it worked without --------- Signed-off-by: Maxime Ellerbach <maxime@ellerbach.net>	2026-05-11 15:36:32 +02:00
Steven Palma	b607c8458e	docs: add policy & compute guide (#3534 ) * docs(policy): contributing a policy guide * docs(training): HW compute guide * chore(docs): add to readme and index * Apply suggestions from code review Co-authored-by: Haoming Song <1847575517@qq.com> Signed-off-by: Steven Palma <imstevenpmwork@ieee.org> * chore(docs): slight improvements * refactor(docs): consolidate add policy docs * chore(style): fix pre-commit --------- Signed-off-by: Steven Palma <imstevenpmwork@ieee.org> Co-authored-by: Haoming Song <1847575517@qq.com>	2026-05-11 15:19:12 +02:00
Jash Shah	9e83510c99	fix(datasets): close file handle on VideoDecoder init failure in cache (#3542 ) If VideoDecoder() raises during initialization, the fsspec file handle was leaked since it was opened via __enter__() but never closed on the exception path. Now explicitly closes the handle before re-raising.	2026-05-10 17:30:37 +02:00
Anthony Shoumikhin	1f7b03f5f2	chore(deps): allow torch 2.11/2.12 and fix autocast deprecation (#3435 ) * chore(deps): allow torch 2.11/2.12 and fix autocast deprecation - Bump torch to >=2.7,<2.13 (was <2.11), torchvision to <0.28 (was <0.26), and torchcodec to <0.13 (was <0.11) to allow installs against the latest stable torch 2.11 and the upcoming 2.12 line. - Replace removed torch.get_autocast_gpu_dtype() with torch.get_autocast_dtype("cuda") in Florence2 and Qwen2.5-VL-MoE FlashAttention paths (the former is removed in 2.11+). - Refresh uv.lock for the new resolution (torch 2.11.0+cu130, torchvision 0.26.0+cu130, torchcodec 0.11.1, full CUDA 13 stack). Verified locally with `uv sync --locked` from a clean .venv and the lerobot test suite (pytest -n 8 --dist=loadfile --timeout=300). Failure set is identical to the pre-bump baseline: 18 pre-existing failures (test_sac_policy, test_pi0_rtc, test_pi05_rtc, test_replay_buffer), 0 new, 0 fixed. AI assistance: this change was authored with Claude Code per AI_POLICY.md. * fix(policies): use device-agnostic autocast dtype lookup Pass query_states.device.type to torch.get_autocast_dtype() instead of hardcoding 'cuda', so the cast matches the active autocast context when running under CPU/MPS/XPU autocast. --------- Co-authored-by: Steven Palma <imstevenpmwork@ieee.org>	2026-05-10 13:05:35 +02:00
Steven Palma	cb8edf17e6	chore(dependencies): update uv.lock (#3475 )	2026-05-10 12:24:22 +02:00
Steven Palma	5699f6cbf4	chore(ci): disable auto-stale (#3550 )	2026-05-10 11:49:31 +02:00
masato-ka	0e6114ac36	fix(train): restrict legacy RA-BC migration to JSON checkpoints only (#3490 ) * fix(train): restrict legacy RA-BC migration to JSON checkpoints only _migrate_legacy_rabc_fields was called for all config files, causing json.load to raise DecodeError when a YAML/TOML config was passed to lerobot-train for a new training run. Guard the block with an .endswith(".json") check so migration only runs when resuming from a JSON checkpoint.	2026-05-08 20:27:01 +02:00
Steven Palma	c8ce413d73	fix(robots): allign lekiwi default with so100 use_degrees (#3531 )	2026-05-07 17:52:34 +02:00
Pepijn	82dffde7fa	fix(ci): speed up multi-task benchmark evals (parallelize + cap VLABench steps) (#3529 ) * fix(ci): run multi-task benchmark evals 5-at-a-time in parallel The eval script supports running tasks concurrently via a ThreadPoolExecutor (env.max_parallel_tasks). Apply it to the four multi-task benchmark CI jobs (RoboTwin, RoboCasa, RoboMME, LIBERO-plus — 8-10 tasks/task_ids each) so they finish in ~2 waves of 5 instead of running sequentially. Single-task jobs (Libero, MetaWorld, RoboCerebra) are unchanged. * fix(ci): cap VLABench smoke eval at 50 steps per task VLABench's default episode_length is 500 steps; with 10 tasks at ~1 it/s the smoke eval took ~80 minutes of rollouts on top of the image build. The eval is a pipeline smoke test (running_success_rate stays at 0% on this short rollout anyway), so we don't need full episodes — cap each task at 50 steps to bring total rollout time down ~10x. * fix(ci): run VLABench tasks 5-at-a-time in parallel The eval script already supports running multiple tasks concurrently via a ThreadPoolExecutor (env.max_parallel_tasks). Set it to 5 so the 10 VLABench tasks finish in ~2 waves instead of running sequentially.	2026-05-07 13:37:16 +02:00
Ville Kuosmanen	eaf0218bc8	feat(policy): use pretrained vision encoder weights by default for diffusion and vqbet (#3202 ) * feat: add pretrained vision encoder weights for diffusion and vqbet * fix test by re-generating artifacts --------- Co-authored-by: Steven Palma <imstevenpmwork@ieee.org>	2026-05-07 12:10:38 +02:00
Pepijn	a0e52d52fe	fix(ci): bump robotwin benchmark image to CUDA 12.6 (#3525 ) The robotwin benchmark Dockerfile still installed cuda-nvcc-12-4 and cuda-cudart-dev-12-4 after #3505 upgraded the base image to CUDA 12.6.3 on Ubuntu 24.04. Those packages aren't available in the ubuntu2404 CUDA repo, so the build failed at apt-get install. Bumping both to -12-6 to match the base image.	2026-05-07 11:11:12 +02:00