lerobot

mirror of https://github.com/huggingface/lerobot.git synced 2026-07-23 01:41:54 +00:00

Author	SHA1	Message	Date
Pepijn	8d982614a6	Merge remote-tracking branch 'origin/main' into codex/model-profiling # Conflicts: # src/lerobot/configs/train.py	2026-04-20 11:32:10 +02:00
Defalt	5c43fa1cce	fix(policies): replace deprecated torch.cuda.amp.autocast with torch.amp.autocast (#3167 ) Co-authored-by: Steven Palma <imstevenpmwork@ieee.org>	2026-04-19 16:25:08 +02:00
k1000dai	3f16d98a9b	episods→episodes (#3410 ) Fixing typo	2026-04-19 12:58:06 +02:00
whats2000	52f508c51c	fix(dataset): cleanup_interrupted_episode wipes image temp dirs (#3405 )	2026-04-19 12:04:24 +02:00
Steven Palma	a8b72d9615	feat(dataset): 2x faster dataloader via parallel decode, uint8 transport, and persistent workers (#3406 ) * feat(dataset): 2xfaster dataloader * fix(dataset): streaming return uint8 decode * fix(tests): adjust normalization step comparison * fix(dataset): with threadexecutor + False default * chore(dataset): make it a config * fix(test): account for uint8 in training path testing	2026-04-19 00:08:22 +02:00
Steven Palma	760220d532	chore(dependencies): update uv.lock (#3365 ) Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>	2026-04-18 22:32:05 +02:00
Shu Jiuhe	a99943ca26	Improve loading performance in `_absolute_to_relative_idx` when remapping indices (#3279 )	2026-04-18 19:28:50 +02:00
Cheng Yin	a9821af61b	fix(record): pass rename_map to make_policy in lerobot-record (#3240 ) * fix(record): pass rename_map to make_policy in lerobot-record Fixes #3181. The rename_map from dataset config was used for preprocessor construction but not passed to make_policy(), causing feature mismatch errors when camera key names differ between dataset and model config. make_policy() already accepts a rename_map parameter and uses it to skip visual feature consistency validation when remapping is active, but lerobot_record.py was not passing it through. * style: fix ruff format for ternary expression --------- Co-authored-by: Steven Palma <imstevenpmwork@ieee.org>	2026-04-17 16:40:08 +02:00
Pepijn	c8df80ae91	Merge remote-tracking branch 'origin/main' into codex/model-profiling	2026-04-17 12:27:11 +01:00
Steven Palma	d4a229444b	fix(ci): not fail when skipped (#3399 )	2026-04-17 12:02:38 +02:00
Pepijn	1ac8e96575	refactor(profiling): shrink lerobot_train.py diff via start()/finalize() Replace the `with profiler or nullcontext():` wrap around the entire training loop with explicit `profiler.start()` / `profiler.finalize()` calls, and tighten `_section(...)` regions in `update_policy` to only wrap the hot calls (forward / backward / optimizer.step). This avoids ~120 lines of pure re-indentation noise while keeping the exact same artifacts on disk and the same public behavior. lerobot_train.py diff vs main: 267 -> 29 changed lines. Made-with: Cursor	2026-04-17 10:59:43 +01:00
Steven Palma	098ebb4d72	feat(ci): send slack notification if latest dependecy test is broken (#3398 )	2026-04-17 11:28:24 +02:00
Pepijn	a6dd28e8b4	fix(profiling): tolerate groot dep-install failure groot's only policy-specific dependency is flash-attn, which has no prebuilt wheel for torch 2.10 and requires nvcc to build from source. The CI image is based on nvidia/cuda:12.4.1-base, which ships the CUDA runtime but not the compiler toolkit, so the source build fails with `/usr/local/cuda/bin/nvcc: No such file or directory`. The repo's own pyproject.toml already carries a TODO acknowledging this: gr00t needs bespoke flash-attn install steps. Treat this as an environmental limitation rather than a regression: dep-install failures for groot are logged via `::warning::` and skip the policy without failing the job. Dep-install failures for any other policy remain fatal, so real regressions still surface. Made-with: Cursor	2026-04-16 21:15:14 +02:00
Pepijn	1842100402	feat(profiling): record forward/backward/optimizer timings The dashboard expects per-phase timings (forward_s, backward_s, optimizer_s) in step_timing_summary.json, but only total_update_s and dataloading_s were collected — leaving every chart except dataloading empty. Add a lightweight TrainingProfiler.section(name) context manager that times a region with torch.cuda.synchronize before and after (so GPU work is captured, not just the kernel-launch latency) and accumulates per-section samples into step_timing_summary.json. Wrap forward, backward (incl. grad clip), and optimizer (incl. zero_grad and scheduler.step) in update_policy with these sections. When profiling is off (profiler=None) the wrappers become no-ops, so training performance is unchanged outside CI. Made-with: Cursor	2026-04-16 20:26:27 +02:00
Pepijn	00e9defb80	fix(profiling): build flash-attn without isolation for groot groot depends on flash-attn, which fails to build in uv's default isolated build env because it doesn't declare torch as a build-time dependency. Torch is a core lerobot dep and is already present in the target venv when groot is synced, so we can safely disable build isolation just for flash-attn. The flag is a no-op for policies that don't pull in flash-attn. Made-with: Cursor	2026-04-16 20:21:58 +02:00
Pepijn	b81eef43c8	fix(profiling): wall_x OOM and xvla rename_map - wall_x: switch to SGD optimizer + explicit scheduler overrides. The 4B-param model casts to bf16 internally, but AdamW's exp_avg/ exp_avg_sq states blow past the 22 GB GPU. Same fix we applied to pi0/pi05/pi0_fast. - xvla: fix rename_map. Dataset (libero_plus) exposes front/wrist image keys; the model expects image/image2. Previous map was direction-reversed and left the batch without any recognized image feature. Made-with: Cursor	2026-04-16 19:49:12 +02:00
Pepijn	d483dd4c4b	feat(profiling): profile groot, xvla, diffusion, wall_x on PRs Add groot, xvla, diffusion and wall_x (wall-oss-flow) to the smoke profiling filter and switch the runner to per-policy dependency resolution. Each policy now gets its own `uv sync --extra <policy>` pass followed by a profiling run, so heavy or conflicting extras (flash-attn, peft, diffusers, etc.) can never block another policy's profiling. A failure in one policy is logged and surfaces a non-zero exit at the end instead of aborting the matrix. Made-with: Cursor	2026-04-16 19:04:27 +02:00
Pepijn	a56423fa33	Merge branch 'main' into codex/model-profiling	2026-04-16 18:58:35 +02:00
Maxime Ellerbach	9bc2df80bb	chore(docs): adding a jupyter notebook that gives you ready-to-paste commands (#3395 ) * chore(docs): adding an example quickstart jupyter notebook that gives you ready-to-paste commands * some fixes in the commands * uv lock * Adding notebook to all Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> Signed-off-by: Maxime Ellerbach <maxime@ellerbach.net> * uv lock again --------- Signed-off-by: Maxime Ellerbach <maxime@ellerbach.net> Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>	2026-04-16 17:53:35 +02:00
Pepijn	da7da741f1	fix(profiling): use SGD for pi0/pi05/pi0_fast and free CUDA cache after deterministic forward Adam optimizer states (exp_avg + exp_avg_sq) require ~16GB extra on top of model params and gradients for 4B parameter models, exceeding the 22GB GPU. SGD has zero optimizer state overhead and profiling only measures forward/backward timing anyway. Also adds torch.cuda.empty_cache() after deterministic forward to release transient memory before the training loop starts. Made-with: Cursor	2026-04-16 16:09:56 +02:00
Pepijn	b1e16783de	refactor: extract profiling into self-contained TrainingProfiler class Move all profiling orchestration out of lerobot_train.py and TrainPipelineConfig into a TrainingProfiler class in profiling_utils.py. - lerobot_train.py: ~74 lines of profiling code reduced to ~7 call sites - TrainPipelineConfig: 10 profile_* fields reduced to 2 (mode + output_dir) - update_policy: reverted to clean main-branch signature (no timing_collector) - TrainingProfiler encapsulates torch profiler, timing collection, deterministic forward artifacts, and all output writing - CI script (run_model_profiling.py) unchanged—it only passes the 2 kept fields Made-with: Cursor	2026-04-16 16:00:49 +02:00
Pepijn	a4544ffea7	fix(profiling): use bf16 dtype and gradient checkpointing for pi0/pi05 Enable --policy.dtype=bfloat16 and --policy.gradient_checkpointing=true for pi0, pi0_fast, and pi05 profiling specs. Combined with use_amp=true, this brings the 4B-param VLA models well within the 22GB GPU budget. Made-with: Cursor	2026-04-16 15:35:25 +02:00
Pepijn	dbe01b0444	fix(profiling): fix pi0 cuBLAS error and pi05 OOM on 22GB GPU - Move cudnn_deterministic to per-spec train_args instead of hardcoding it for all models. cuBLAS deterministic mode triggers internal errors on Gemma-based models (pi0, pi05) during backward pass. - Enable use_amp=true for pi0, pi0_fast, and pi05 to reduce memory footprint from fp32 (~16GB weights alone) to bf16, fitting within 22GB GPU budget with room for activations and gradients. - Small models (act, diffusion, multi_task_dit) still use deterministic mode for reproducible profiling results. Made-with: Cursor	2026-04-16 15:34:17 +02:00
Pepijn	e16a95a78e	refactor(profiling): remove cProfile, keep torch profiler only Remove cProfile wrapping from the training loop and profiling utilities. The torch profiler already captures fine-grained timing and operator breakdowns; cProfile added redundant overhead without actionable insight for GPU-bound models. - Remove render_cprofile_summary, run_with_cprofile from profiling_utils - Replace cProfile-wrapped calls in lerobot_train with direct calls - Remove cprofile_summaries from artifact index in run_model_profiling - Update tests to match Made-with: Cursor	2026-04-16 15:34:17 +02:00
Pepijn	4137b5785d	fix(profiling): align libero smoke specs with pretrained policies	2026-04-16 15:11:54 +02:00
Pepijn	8ece10e484	feat(ci): profile more models in pr smoke runs	2026-04-16 14:49:37 +02:00
Pepijn	ddeb216ab9	fix(ci): skip hub publish for pr profiling runs	2026-04-16 14:38:43 +02:00
Pepijn	d46d67f75d	fix(profiling): forward GIT_REF + PR_NUMBER into Docker container The previous commit moved these expressions from inline shell expansion to job-level env: vars, but the profiling script runs inside a Docker container. Job-level env vars are only visible in the runner, not inside the container — they need explicit -e flags on the docker run command (same pattern as HOST_GIT_COMMIT which was already forwarded). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-16 13:38:13 +02:00
Pepijn	b746cd3c61	fix(profiling): sort import + move expressions to env vars for zizmor Pre-commit Quality gate flagged two issues: 1. ruff/isort: `from numbers import Real` must sort after `from collections.abc import Callable` (stdlib alphabetical order). 2. zizmor (high): `github.head_ref`, `github.ref_name`, `github.event.inputs.git_ref`, and `github.event.pull_request.head.sha` were expanded directly in `run:` shell blocks, which zizmor flags as attacker-controllable. Move all four into job-level `env:` vars (GIT_REF, PR_NUMBER, HOST_GIT_COMMIT) so the shell only sees env-var references — the same pattern the workflow already uses for PROFILE_MODE, POLICY_FILTER, etc. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-16 13:30:13 +02:00
Pepijn	6d1a5fca02	fix(profiling): keep ci green when hub publish is unauthorized	2026-04-16 13:07:30 +02:00
Pepijn	8d7099cd7d	fix(profiling): publish preview runs via hf dataset prs	2026-04-16 12:50:57 +02:00
Pepijn	516f39685a	fix(profiling): skip dataset creation on publish	2026-04-16 12:09:03 +02:00
Pepijn	b27e838376	fix(profiling): publish preview rows to existing dataset	2026-04-16 11:54:35 +02:00
Pepijn	40470648d1	feat(profiling): publish preview runs for dashboard debugging	2026-04-16 10:54:34 +02:00
Pepijn	25e5062b2c	fix(profiling): read generic device timings from profiler	2026-04-16 10:29:01 +02:00
Pepijn	35e3b28da1	fix(profiling): normalize timing metrics before export	2026-04-16 10:11:14 +02:00
Pepijn	ed8a98dda6	fix(profiling): preserve policy mode for deterministic forward	2026-04-16 09:50:29 +02:00
Pepijn	9dc38d9993	fix(ci): isolate torch cache in profiling job	2026-04-16 09:32:16 +02:00
Pepijn	3922f81791	fix(ci): set HF_LEROBOT_HOME in profiling job	2026-04-15 23:35:27 +02:00
Pepijn	28e8483297	fix(ci): disable policy hub push in profiling runs	2026-04-15 23:02:28 +02:00
Pepijn	e1b22ed1c4	fix(ci): set torchinductor cache dir in profiling job	2026-04-15 22:55:31 +02:00
Pepijn	f2d0f04dd0	fix(ci): isolate profiling container home dirs	2026-04-15 22:51:22 +02:00
Pepijn	3ea722c6c0	fix(ci): run profiling container as runner user	2026-04-15 22:47:29 +02:00
Pepijn	48660e7a7c	fix(ci): avoid host shell expansion in policy error	2026-04-15 22:42:34 +02:00
Pepijn	c94fe868c9	fix(ci): install only profiling policy extras	2026-04-15 22:38:37 +02:00
Pepijn	d4f27cfb6e	fix(ci): restore docker env line continuation	2026-04-15 22:33:14 +02:00
Pepijn	1a2aec1b04	feat(profiling): add weekly model profiling	2026-04-15 22:31:44 +02:00
Remy	bd74f6733d	chore: bump doc-builder SHA for PR upload workflow (#3386 )	2026-04-15 12:15:24 +02:00
Steven Palma	6f4a96333e	chore(docs): update contributing (#3387 )	2026-04-15 11:02:37 +02:00
Steven Palma	9021d2d240	refactor(imports): enforce guard pattern (#3382 ) * refactor(imports): enforce guard pattern * fix(tests): skip reachy2 if not installed * Address review feedback	2026-04-14 22:54:05 +02:00

1 2 3 4 5 ...

1435 Commits