mirror of
https://github.com/huggingface/lerobot.git
synced 2026-07-24 18:26:11 +00:00
cec8ee0be6
Steerable annotation pipeline (lerobot-annotate) that populates the language_persistent and language_events columns introduced in PR 1 (#3467) directly into data/chunk-*/file-*.parquet. This is PR 2 of the three-PR plan: PR 1 (Add extensive language support #3467): schema + DSL + rendering, base of this PR PR 2 (this PR): annotation pipeline writing into PR 1's columns PR 3: model with language prediction and runtime A VLM (Qwen-VL family, served on vLLM) watches each episode's video and emits grounded language annotations: subtasks, plans, memory, task rephrasings, interjections + speech, and per-camera VQA. The pipeline is built for production annotation at scale — single-camera grounding, embedded-frame inputs, a describe-then-segment grounding flow, and a deterministic full-episode coverage guarantee — informed by Scale's dense-captioning findings (representation > sampling, rules > reasoning, model capacity is the biggest lever, two-pass systems compound errors)
33 lines
1.3 KiB
Plaintext
33 lines
1.3 KiB
Plaintext
You are generating a frame-grounded visual question/answer pair for
|
|
chain-of-thought training. Reference: ECoT (Zawalski 2024) and Steerable
|
|
Policies — both train policies on grounded features such as bounding box
|
|
pixel coordinates, keypoints, counts, attributes, and spatial relations.
|
|
|
|
The frame shows a robot working on: "{episode_task}".
|
|
|
|
Question types and the EXACT answer JSON shape required for each:
|
|
|
|
bbox => {{"detections": [{{"label": "<obj>", "bbox_format": "xyxy",
|
|
"bbox": [x1, y1, x2, y2]}}, ...]}}
|
|
bbox is in pixel coordinates (x_min, y_min, x_max, y_max).
|
|
ECoT example: "a white cup [124, 25, 176, 113]".
|
|
|
|
keypoint => {{"label": "<point>", "point_format": "xy",
|
|
"point": [x, y]}}
|
|
|
|
count => {{"label": "<obj>", "count": <int>,
|
|
"note": "<optional short note>"}}
|
|
|
|
attribute => {{"label": "<obj>", "attribute": "<color|shape|state|...>",
|
|
"value": "<observed value>"}}
|
|
|
|
spatial => {{"subject": "<obj>", "relation": "<left_of|right_of|on|in|"
|
|
"above|below|near>", "object": "<obj>"}}
|
|
|
|
Generate a question of type "{question_type}". Output strictly valid JSON:
|
|
|
|
{{
|
|
"question": "<short, frame-grounded question>",
|
|
"answer": <object whose shape matches the schema above>
|
|
}}
|