lerobot

mirror of https://github.com/huggingface/lerobot.git synced 2026-07-08 18:41:54 +00:00

Files

T

Pepijn c026aed8f8 feat(pi052): train VQA spatial answers in PaliGemma <loc> format

Spatial VQA answers (bbox / keypoint) were trained as pixel-coordinate
JSON, which fights PaliGemma's detection prior and leaks <loc>-token
salad at inference. Convert them to PaliGemma's native <locNNNN>
vocabulary instead so the LM head reuses that prior.

Training side (text_processor_pi052.py): a target turn whose content
parses as a bbox/keypoint answer is rewritten to <loc> text, using the
camera frame's native (H, W) from the observation and the preceding
image block. Non-spatial answers, subtask/memory targets and SmolVLA2
keep their JSON form — the dataset stays backbone-agnostic.

Runtime side (smolvla2/inference/vqa.py): parse_vqa_answer detects
<loc> answers (2 locs -> keypoint, 4 -> bbox), returning normalized
[0,1] coords with a normalized flag; draw_vqa_overlay denormalizes
against the chosen camera frame's pixel size.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

2026-05-19 20:23:46 +02:00

groot

feat(dependencies): minimal default tag install (#3362 )

2026-04-12 20:03:04 +02:00

hilserl

feat(dependencies): minimal default tag install (#3362 )

2026-04-12 20:03:04 +02:00

multi_task_dit

fix(test): add missing device placement in multi-task DiT tests (#3349 )

2026-04-14 12:25:29 +02:00

pi0_fast

chore(dependecies): untangle dependecies across internal modules (#3149 )