feat(pi052): train VQA spatial answers in PaliGemma <loc> format

mirror of https://github.com/huggingface/lerobot.git synced 2026-07-12 20:41:58 +00:00

Spatial VQA answers (bbox / keypoint) were trained as pixel-coordinate
JSON, which fights PaliGemma's detection prior and leaks <loc>-token
salad at inference. Convert them to PaliGemma's native <locNNNN>
vocabulary instead so the LM head reuses that prior.

Training side (text_processor_pi052.py): a target turn whose content
parses as a bbox/keypoint answer is rewritten to <loc> text, using the
camera frame's native (H, W) from the observation and the preceding
image block. Non-spatial answers, subtask/memory targets and SmolVLA2
keep their JSON form — the dataset stays backbone-agnostic.

Runtime side (smolvla2/inference/vqa.py): parse_vqa_answer detects
<loc> answers (2 locs -> keypoint, 4 -> bbox), returning normalized
[0,1] coords with a normalized flag; draw_vqa_overlay denormalizes
against the chosen camera frame's pixel size.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

This commit is contained in:

Pepijn

2026-05-19 20:23:46 +02:00

parent e425dfd624

commit c026aed8f8

5 changed files with 1917 additions and 141 deletions

uv.lock

Generated

+1528 -137

View File

File diff suppressed because it is too large Load Diff