Files
lerobot/.github/workflows
Pepijn 893b10b825 ci(vlabench): smoke-eval 10 tasks instead of 1
Broader coverage on the VLABench benchmark CI job: bump the smoke eval
from 1 task to 10 (one episode each), all drawn from PRIMITIVE_TASKS.

Tasks now run: select_fruit, select_toy, select_book, select_painting,
select_drink, select_ingredient, select_mahjong, select_poker,
add_condiment, insert_flower.

Bumped the parse_eval_metrics.py `--task` label to the full comma
list so the metrics artifact reflects what was actually run.
`parse_eval_metrics.py` already reads `overall` for multi-task runs,
so no parser change is needed.

Made-with: Cursor
2026-04-17 13:40:42 +01:00
..