mirror of
https://github.com/huggingface/lerobot.git
synced 2026-07-26 11:16:00 +00:00
893b10b825
Broader coverage on the VLABench benchmark CI job: bump the smoke eval from 1 task to 10 (one episode each), all drawn from PRIMITIVE_TASKS. Tasks now run: select_fruit, select_toy, select_book, select_painting, select_drink, select_ingredient, select_mahjong, select_poker, add_condiment, insert_flower. Bumped the parse_eval_metrics.py `--task` label to the full comma list so the metrics artifact reflects what was actually run. `parse_eval_metrics.py` already reads `overall` for multi-task runs, so no parser change is needed. Made-with: Cursor