mirror of
https://github.com/huggingface/lerobot.git
synced 2026-07-30 13:09:40 +00:00
279c6c7af3
* feat(annotate): WGO-tuned subtask prompt (atomic completed-events + duration prior)
Rework the plan-module subtask segmentation prompt toward the WGO-Bench
atomic annotation protocol: segment by completed world-state changes
(grasp/place/open/close/pour/insert), fold approach+retreat into their
event, keep separate events separate, and add a 2-10s duration prior.
Drops the pi0.7 "fewer larger composites preferred" bias that drove
under-segmentation on the benchmark. Output JSON shape unchanged.
Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(annotate): seeded-relabeling second pass for subtasks
Add an opt-in relabel pass (plan.subtask_seeded_relabel) that, after
segmentation, re-labels each span using previous/current/next segment
contact sheets and the seed label as a strong prior, minimally correcting
it. Mirrors macrodata's best end-to-end labeling step. Boundaries are
untouched; one extra VLM call per span. Off by default.
Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(annotate): robust OpenAI-compat client for hosted VLMs
Guard against a choice with no message (safety filter or a thinking model
that spends its whole budget before emitting content) so one empty reply
no longer crashes the whole annotation run; treat it as an empty response
and let the existing JSON-retry path handle it.
Add an optional `reasoning_effort` knob on VlmConfig, forwarded to the
server when set, to cap a thinking model's reasoning (needed for Gemini
via its OpenAI-compatible endpoint).
Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(annotate): legible tile-scaled timestamp on contact sheets
The burned-in timestamp used the ~10px bitmap default font, which blurs
once the model downsamples a full contact sheet into 768px tiles, so the
VLM can no longer read the exact source time a boundary depends on. Scale
the timestamp to the tile height (with a graceful fallback on older
Pillow) so the visual time cue stays readable at sheet resolution.
Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(annotate): lean GEPA-aligned subtask segmentation prompt
Replace the verbose, label-heavy segmentation prompt with a lean
adaptation of the blog's GEPA-found completed_events_duration_prior
recipe: focus on completed manipulation events, explicit no-split /
no-merge rules, a 2-10s duration prior, and an instruction to prioritize
temporally correct boundaries over label wording. The previous prompt
over-weighted label guidance, which traded away boundary precision.
Co-authored-by: Cursor <cursoragent@cursor.com>
* revert: restore original subtask segmentation prompt
The lean GEPA-aligned paraphrase (dd4b0110d) regressed Gemini on the
30-ep subset: Seg F1 0.259 -> 0.189 and E2E 0.184 -> 0.135, driven by
worse under-segmentation (224 -> 188 preds). The blog's 0.306 came from
the actual GEPA-search artifact, which a hand paraphrase does not
reproduce. Restore the original prompt, which remains our best config.
Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(annotate): env-var override for prompt templates
Allow LEROBOT_PROMPT_OVERRIDE_<name> to supersede the packaged prompt
file at load time. Enables prompt search (GEPA) to inject candidate
segmentation prompts into a remote annotate job via an env secret,
without committing a branch per candidate.
Co-authored-by: Cursor <cursoragent@cursor.com>
* docs(annotate): genericize hosted-VLM comments (no model name)
Co-authored-by: Cursor <cursoragent@cursor.com>
* docs(annotate): document seeded-relabel and reasoning_effort flags
Co-authored-by: Cursor <cursoragent@cursor.com>
* test(annotate): update subtask-prompt marker to match WGO-tuned prompt
The three plan-module tests keyed the canned VLM responder on the
literal 'atomic subtasks', which the WGO-tuned segmentation prompt no
longer contains (it now segments 'COMPLETED manipulation events'). Point
the fixture markers at the current wording so the subtask call is matched
again.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
69 lines
3.3 KiB
Plaintext
69 lines
3.3 KiB
Plaintext
You are annotating a teleoperated robot demonstration shown as
|
|
timestamped contact sheets (each tile has its time in seconds burned
|
|
into the top-left corner). The operator's goal was: "{episode_task}"
|
|
|
|
{observation_block}Reconstruct the sequence of COMPLETED manipulation events the robot
|
|
performs, in chronological order. Output one segment per event with a
|
|
[start, end] time in seconds and a short action label.
|
|
|
|
GROUNDING — read first, it overrides everything below:
|
|
- Label ONLY events you can SEE in the frames. The instruction is the
|
|
goal; the VIDEO is the ground truth for what actually happened.
|
|
- Do NOT invent, anticipate, or pad steps that are not shown.
|
|
|
|
Granularity — segment by completed events, not by motion:
|
|
- Start a NEW segment whenever the world state changes: an object is
|
|
grasped, lifted, transported, placed, or released; a held object
|
|
changes; a drawer/door/lid/container opens or closes; contents move
|
|
between containers (poured); a tool starts or stops acting on a
|
|
surface. Watch the gripper open/close transitions — they usually mark
|
|
boundaries.
|
|
- Do NOT split approach, reach, grasp adjustment, small repositioning,
|
|
hesitation, or retreat into their own segments. Fold each into the
|
|
event it belongs to (the approach is part of the pick; the retreat is
|
|
part of the place).
|
|
- Do NOT merge separate completed events. Each distinct pick, place,
|
|
open, close, pour, push, wipe, or insert is its own segment, even when
|
|
they repeat on different objects or locations.
|
|
- Most segments last 2-10 seconds. Shorter segments are okay ONLY for
|
|
fast pick / place / open / close / release events. Never emit a
|
|
segment shorter than {min_subtask_seconds} seconds; merge a too-short
|
|
candidate into its neighbour instead.
|
|
- Skip idle time, pure camera motion, and tiny hand jitter.
|
|
|
|
Labels — short imperative phrases:
|
|
- One concise command naming the action and the manipulated object, e.g.
|
|
"pick up the red cup", "put the cup on the shelf", "open the top
|
|
drawer", "pour water into the glass", "insert the plug into the
|
|
socket".
|
|
- Include source, destination, side, direction, or the final
|
|
open/closed state when it is visible and central to the event.
|
|
- Prefer these verbs (extend only when none fits): pick up, put, place,
|
|
push, pull, turn, press, open, close, pour, insert, wipe, stack.
|
|
Disambiguate by what you SEE:
|
|
* STACK vs PUT: object placed ON TOP OF another object -> "stack".
|
|
* INSERT vs PUT: object pushed INTO a fitted slot/hole/socket -> "insert".
|
|
* PICK UP vs PUT (direction): gripper CLOSES and object moves WITH
|
|
the hand -> "pick up"; gripper OPENS and object stays -> "put".
|
|
* POUR vs PUT: source is tilted and contents flow -> "pour".
|
|
- Use the exact object nouns implied by the task; stay consistent across
|
|
the episode (don't switch "cube" to "block").
|
|
- Write imperative commands, never third person ("the robot ..."), and
|
|
drop articles/adverbs.
|
|
|
|
Timing:
|
|
- Use the burned-in timestamps to set start and end. Boundaries should
|
|
land on or near a printed time, and every [start, end] must lie within
|
|
[0.0, {episode_duration}] seconds, be non-overlapping, and cover the
|
|
episode in order.
|
|
- Emit at most {max_steps} segments.
|
|
|
|
Output strictly valid JSON of shape:
|
|
|
|
{{
|
|
"subtasks": [
|
|
{{"text": "<short imperative action label>", "start": <float>, "end": <float>}},
|
|
...
|
|
]
|
|
}}
|