mirror of
https://github.com/huggingface/lerobot.git
synced 2026-07-24 10:16:09 +00:00
cec8ee0be6
Steerable annotation pipeline (lerobot-annotate) that populates the language_persistent and language_events columns introduced in PR 1 (#3467) directly into data/chunk-*/file-*.parquet. This is PR 2 of the three-PR plan: PR 1 (Add extensive language support #3467): schema + DSL + rendering, base of this PR PR 2 (this PR): annotation pipeline writing into PR 1's columns PR 3: model with language prediction and runtime A VLM (Qwen-VL family, served on vLLM) watches each episode's video and emits grounded language annotations: subtasks, plans, memory, task rephrasings, interjections + speech, and per-camera VQA. The pipeline is built for production annotation at scale — single-camera grounding, embedded-frame inputs, a describe-then-segment grounding flow, and a deterministic full-episode coverage guarantee — informed by Scale's dense-captioning findings (representation > sampling, rules > reasoning, model capacity is the biggest lever, two-pass systems compound errors)
113 lines
5.9 KiB
Plaintext
113 lines
5.9 KiB
Plaintext
You are labeling a teleoperated robot demonstration.
|
|
|
|
The user originally asked: "{episode_task}"
|
|
|
|
You are shown the entire demonstration as a single video. Watch the
|
|
whole clip, then segment it into a list of consecutive atomic subtasks
|
|
the robot performs.
|
|
|
|
{observation_block}GROUNDING — read this first, it overrides everything below:
|
|
- Label ONLY what the robot actually does in the video. Every subtask
|
|
you emit must correspond to motion you can SEE in specific frames.
|
|
- Do NOT invent, anticipate, or pad. If the robot only does one thing
|
|
(e.g. it just navigates to a location and the clip ends), emit
|
|
EXACTLY ONE subtask. Many demonstrations are a single atomic skill.
|
|
- ``max_steps`` below is a hard CEILING, not a target. Emitting fewer
|
|
subtasks than the ceiling is not just allowed, it is expected for
|
|
short / atomic demonstrations. One correct subtask is far better
|
|
than several invented ones.
|
|
- If the video does not clearly show the action implied by the task,
|
|
describe what you actually see — do NOT fabricate the task's steps
|
|
from the instruction text. The instruction tells you the goal; the
|
|
VIDEO is the ground truth for what happened.
|
|
|
|
Authoring rules — Hi Robot atom granularity, pi0.7-style short prompts:
|
|
|
|
- Each subtask = one COMPOSITE atomic skill the low-level policy can
|
|
execute end-to-end. A "skill" bundles its own approach motion with
|
|
its terminal action — do NOT split the approach off as its own
|
|
subtask. The whole-arm policy already learns to reach as part of
|
|
every manipulation primitive.
|
|
- Write each subtask as an IMPERATIVE COMMAND, starting with one of
|
|
these verbs (extend only when none fits):
|
|
pick up <obj> — approach + grasp + lift in one subtask
|
|
put <obj> on/in <loc> — transport + release in one subtask
|
|
place <obj> on/in <loc> — synonym of "put"; pick one and stay consistent
|
|
push <obj> — contact + linear shove
|
|
pull <obj> — contact + linear retract
|
|
turn <knob/dial/handle> — rotary actuation
|
|
press <button> — single-press contact
|
|
open <drawer/door/lid> — full open motion
|
|
close <drawer/door/lid> — full close motion
|
|
pour <src> into <dst> — tilt + flow
|
|
insert <obj> into <slot>— alignment + push-fit
|
|
go to <loc> — ONLY when no grasp / actuation follows
|
|
(e.g. a pure relocation between phases).
|
|
If the next subtask grasps something at
|
|
that location, drop "go to ..." and just
|
|
write "pick up ..." instead.
|
|
- Forbidden ultra-fine splits — the VLM is NOT allowed to emit these
|
|
as standalone subtasks; fold them into the parent composite:
|
|
"move to X" → fold into "pick up X" (or whatever follows)
|
|
"reach for X" → fold into "pick up X"
|
|
"grasp X" → fold into "pick up X"
|
|
"lift X" → fold into "pick up X" (or "put X on Y" if it's
|
|
the transport phase of a place)
|
|
"release X" → fold into "put X on Y" (or "place X in Y")
|
|
- Keep it SHORT — a verb phrase, not a sentence. Drop articles
|
|
("the", "a") and adverbs ("carefully", "slowly"). Add a "how"
|
|
detail (which hand, which grasp point) ONLY when it is needed to
|
|
disambiguate. Every subtask must begin with one of the verbs
|
|
above (no leading nouns, no "then", no "first").
|
|
- NEVER use third person. Never write "the robot", "the arm", "the
|
|
gripper moves", "it picks up" — the robot is implied. Command it,
|
|
do not describe it.
|
|
- Use the exact object nouns from the task above. If the task says
|
|
"cube", every subtask says "cube" — never switch to "block". If it
|
|
says "box", never switch to "bin"/"container". Keep vocabulary
|
|
consistent across the whole episode.
|
|
- Good: "pick up blue cube", "put blue cube in box", "open drawer",
|
|
"turn red knob", "press start button", "go to sink".
|
|
- Bad: "move to blue cube" (approach as its own subtask — forbidden,
|
|
must be folded into "pick up blue cube"); "the robot arm moves
|
|
towards the blue cube" (third person, too long); "carefully pick
|
|
up the cube" (adverb, article); "release the yellow block"
|
|
("block" when the task said "cube", and "release" must be folded
|
|
into a "put"/"place" subtask).
|
|
- Subtasks are non-overlapping and cover the full episode in order.
|
|
Choose the cut points yourself based on what you see in the video
|
|
(gripper open/close events, contact, regrasps, transitions).
|
|
- Each subtask spans at least {min_subtask_seconds} seconds. If a
|
|
candidate span would be shorter, merge it into its neighbour
|
|
rather than emitting it.
|
|
- Do not exceed {max_steps} subtasks total. Fewer, larger composites
|
|
are preferred over many micro-steps.
|
|
- Every subtask's [start_time, end_time] must lie within
|
|
[0.0, {episode_duration}] seconds.
|
|
|
|
SPECIAL CASES — verb disambiguation (each rule is narrowly visual and
|
|
fires ONLY on the spatial situation it names; it must not change how you
|
|
label any other situation):
|
|
- STACK vs PUT: if an object is placed ON TOP OF another specific object
|
|
(not on a flat table / shelf / counter), use "stack ... on ...", not
|
|
"put". "stack blue book on green book", NOT "put blue book on table".
|
|
- INSERT vs PUT: if an object goes INTO a fitted slot / hole / socket /
|
|
receptacle (push-fit), use "insert ... into ...", not "put".
|
|
- RETRIEVE/PICK-UP vs PUT (direction): watch the gripper. If it CLOSES
|
|
on the object and the object moves WITH the hand, it is "pick up" /
|
|
"retrieve" (object leaves its location). If the gripper OPENS and the
|
|
object stays where the hand left it, it is "put" / "place" (object
|
|
arrives at a location). Decide by which way the object moves, not by
|
|
where the hand ends up.
|
|
- POUR vs PUT: only use "pour" when the source is tilted and contents
|
|
flow out; moving a full container without tilting is "put"/"place".
|
|
|
|
Output strictly valid JSON of shape:
|
|
|
|
{{
|
|
"subtasks": [
|
|
{{"text": "<short imperative verb phrase>", "start": <float>, "end": <float>}},
|
|
...
|
|
]
|
|
}}
|