mirror of
https://github.com/huggingface/lerobot.git
synced 2026-07-24 02:06:15 +00:00
cec8ee0be6
Steerable annotation pipeline (lerobot-annotate) that populates the language_persistent and language_events columns introduced in PR 1 (#3467) directly into data/chunk-*/file-*.parquet. This is PR 2 of the three-PR plan: PR 1 (Add extensive language support #3467): schema + DSL + rendering, base of this PR PR 2 (this PR): annotation pipeline writing into PR 1's columns PR 3: model with language prediction and runtime A VLM (Qwen-VL family, served on vLLM) watches each episode's video and emits grounded language annotations: subtasks, plans, memory, task rephrasings, interjections + speech, and per-camera VQA. The pipeline is built for production annotation at scale — single-camera grounding, embedded-frame inputs, a describe-then-segment grounding flow, and a deterministic full-episode coverage guarantee — informed by Scale's dense-captioning findings (representation > sampling, rules > reasoning, model capacity is the biggest lever, two-pass systems compound errors)
47 lines
2.0 KiB
Plaintext
47 lines
2.0 KiB
Plaintext
You are generating training data for a Hi Robot-style hierarchical
|
|
robot policy. The robot in this demonstration has ALREADY executed
|
|
every step shown in the video — we cannot retroactively change the
|
|
action stream. To keep training data consistent with the video, the
|
|
"interjection" must align with what the robot is *about to do next* in
|
|
the demonstration, framed as a natural mid-task user request.
|
|
|
|
The episode's overall task: "{episode_task}".
|
|
|
|
The images above show roughly {window_seconds:.1f} seconds straddling a
|
|
subtask boundary in the demonstration:
|
|
|
|
- Subtask the robot just finished: "{prev_subtask}"
|
|
- Subtask the robot is about to start: "{next_subtask}"
|
|
- Time into episode: {timestamp:.2f}s
|
|
|
|
Write ONE compact interjection the user would naturally say at this
|
|
moment to prompt / confirm / encourage the robot to do "{next_subtask}".
|
|
Keep it like a mid-task coaching cue, not a full instruction paragraph.
|
|
Also write the robot's compact verbal acknowledgement.
|
|
|
|
Hard rules:
|
|
|
|
- The interjection MUST be consistent with the next subtask. The user
|
|
cannot ask for something different from what the robot then does in
|
|
the video. If you're tempted to say "actually skip X" or "do Y
|
|
instead", DO NOT — those would contradict the demonstration.
|
|
- The interjection must reference an object, location, or action that
|
|
is plausible given the visible scene and the next subtask text.
|
|
- One short phrase or sentence each. Conversational, not robotic.
|
|
- Prefer direct cues: "{next_subtask}, please."; "Now {next_subtask}."
|
|
- Keep robot speech very short: "OK.", "On it.", "Doing that."
|
|
|
|
Style examples (vary the phrasing — don't reuse these verbatim):
|
|
- "Now go ahead and {next_subtask}."
|
|
- "Great, can you {next_subtask} next?"
|
|
- "{next_subtask}, please."
|
|
- "Before you continue, please {next_subtask}."
|
|
- "Looking good — {next_subtask} now."
|
|
- "Okay, {next_subtask}."
|
|
|
|
Output strictly valid JSON:
|
|
{{
|
|
"interjection": "<short cue from the user, asking for the next subtask>",
|
|
"speech": "<short robot acknowledgement>"
|
|
}}
|