mirror of
https://github.com/huggingface/lerobot.git
synced 2026-07-29 04:36:04 +00:00
3dd19d043e
* feat(depth): add depth quantization helpers and tests
* feat(video): add ffv1 to supported codecs
* feat(depth): persist depth metadata
* feat(depth): extend quantization tools to better fit the encoding/decoding pipeline
* feat(depth): plumb DepthEncoderConfig through LeRobotDataset and DatasetWriter
* feat(depth): wire StreamingVideoEncoder + writer to depth encoder
* feat(depth): wire DatasetReader to decode_depth_frames
* feat(cameras/realsense): expose async depth in metric meters
* feat(features): route 2D camera shapes to observation.depth.<key>
* feat(robots/so_follower): emit + populate depth keys when use_depth
* feat(record): plumb DepthEncoderConfig through lerobot-record
* feat(viz): render depth observations as rr.DepthImage in Viridis
* feat(depth maps writer): adding support for raw depth maps recording with image writer
* chore(format): format code
* feat(depth shape): ensuring depth maps shape is always including the channel
* feat(is_depth): simplifying is_depth nested name + legacy support
* fix(stop_event): fixing stop_event race condition in camera classes
* fix(plumbing): fixing missing parts in the depth maps pipeline
* chore(typos): fixing typos
* test(fix): fixing exisiting tests to still work with latest features
* tests(depth): adding new tests for depth integration validation
* feat(pix_fmt channels): use PyAv to check get pixel formats number of channels
* feat(refactor): refactor DepthEncoderConfig quantization pipeline, so that the methods do not live in the config class. Add pixel format - channels validation.Move the default pixel format for depth in the config file.
* fix(pre-commit): fixing mutable defautl value
* fix(info): fixing info metadata update when is_depth_map was set
* tests(typos): fixing typos in tests
* fix(realsense): fixing typo in realsense serial number
* fix(normalization): restricting 255 normalization to non depth/uint8 images only
* fix(typo): fixing typo
* fix(TIFF): add missing quantization and cleanup for TIFF files
* feat(batched dequantization): optimizing dequantize_depth for torch based batched dequantization
* feat(tools): adding depth support in LeRobotDataset edition tools
* test(aggregate): extending aggregation tests to depth frames
* test(cleaning): cleaning up tests
* fix(from_video_info): fixing early validation issue in from_video_info
* fix(typo): fixing typo
* fix(is_depth): adding missing doctrings and is_depth arguments in video decoding functions
Co-authored-by: Wensi (Vince) Ai <59036629+wensi-ai@users.noreply.github.com>
* fix(depth units): fixing depth units output for the realsense cameras
* feat(output unit): adding support for output unit specification at dataset reading/training time
Co-authored-by: Wensi (Vince) Ai <59036629+wensi-ai@users.noreply.github.com>
* test(depth): cleaning up depth tests
* test(depth encoding): updating and cleaning video/depth encoding tests
* chore(format): formatting code
* docs(depth): improving depth maps docs
* test(fix): fixing depth tests
* test(dataset tools): adding missing tests for new dataset edition tools features
* chore(format): formatting code
* fix(pyav check): fixing PyAV option validation for integer codec options by normalizing
numeric values before calling `is_integer()`
Co-authored-by: Wensi (Vince) Ai <59036629+wensi-ai@users.noreply.github.com>
* docs(mermaid): fixing mermaid diagram
* fix(rebase): rebase follow up corrections
* feat(dataset tools): adding missing docstrings and features for depth fill support in dataset edition tools
* docs(docstring): updating docstrings
* docs(dataset tools): updating docs
* fix(save images): fixing image saving in dataset tools
* fix(update video info): fixing update video info logic to match the recording and editing use cases
* test(reencode): fixing reencoding monkeypatch
* fix(review): add Claude review
* chore(format): format code
* fix(update video info): ditching the differentiated approahces for video info update - video info are always updated unless for preserved keys.
* chore(rebase): fixing rebase merge conflicts
* test(visualization): fixing visualization tests
* feat(docstrings): adding explicit docstring for encoding parameters. Docstrigns will now show up as description in the CLI --help.
* feat(mm as default): adding a global DEFAULT_DEPTH_UNIT variable setting mm as default depth unit
* fix(RGB <-> camera): renaming camera_encoder to rgb_encoder for clarity
* chore(TODO): removing deprecated TODO
* doc(write_u16_plane): improving docstrings for write_u16_plane
* feat(units): adding constants for depth frames units (m and mm)
* fix(spam): replacing spamming warning but a debug log
* feat(leagcy metadata): adding automatic metadata update for legacy 'video.is_depth_map' feature
* fix(copy&reindex): fixing metadat reshaping for single channel frames
* fix(ImageNet): excluding dpeth frames from ImageNet stats
* fix(PyAV container seek): fixing initial PyAV container seek to be robust againsy codec choice
* feat(lerobot-dataset-viz): adding support for depth in lerobot-dataset-viz
* fix(compress): removing rerun compression for DepthImages
* fix(signle channel squeeze): fixing single channel squeezing
* chore(format): format code
* fix(streaming): adding support for dequantization in streaming_dataset.py
* refactor(read depth): factorizing depth reading methods for realsense camera and adding support for depth-only usage
* chore(renaming): fixing missed RGBEncoderConfig renamings
* docs(renaming): reflecting renamings in a clearer way in the docs
* chore(annotation): excluding depth from the annotation pipeline
* feat(robots): adding depth support in compatible follower robots
* feat(LeSadKiwi): excluding LeKiwi from depth support (for now)
* chore(fail): removing misplaced file
* chore(fail): removing misplaced file
* fix(remove ffv1): removing ffv1 as it does not support MP4
* docs(cheat sheet): adding depth and video encoding to the cheat sheet
* fix(lossless): tuning depth encoding parameters for lossless depth storage
* test(fix): fixing failing tests
* depth(ZMQ): excluding ZMQ from depth support
* Revert "depth(ZMQ): excluding ZMQ from depth support"
This reverts commit b95cf4e4c2.
* fix(image transforms): excluding depth frames from images transforms
* fix(typo): typo
* fix(stats): fixing stats computation for depth frames
* fix(TIFF vs. pytorch): adding an extra uint16 to float32 conversion for depth maps stored as raw TIFF images
* fix(typos): fixing typos
* test(dtype): fixing stats computation typing tests
---------
Signed-off-by: Steven Palma <imstevenpmwork@ieee.org>
Co-authored-by: Wensi (Vince) Ai <59036629+wensi-ai@users.noreply.github.com>
Co-authored-by: Steven Palma <imstevenpmwork@ieee.org>
Co-authored-by: Wensi Ai <wsai@stanford.edu>
262 lines
9.6 KiB
Python
262 lines
9.6 KiB
Python
#!/usr/bin/env python
|
|
|
|
# Copyright 2024 The HuggingFace Inc. team. All rights reserved.
|
|
#
|
|
# Licensed under the Apache License, Version 2.0 (the "License");
|
|
# you may not use this file except in compliance with the License.
|
|
# You may obtain a copy of the License at
|
|
#
|
|
# http://www.apache.org/licenses/LICENSE-2.0
|
|
#
|
|
# Unless required by applicable law or agreed to in writing, software
|
|
# distributed under the License is distributed on an "AS IS" BASIS,
|
|
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
# See the License for the specific language governing permissions and
|
|
# limitations under the License.
|
|
import logging
|
|
import multiprocessing
|
|
import queue
|
|
import threading
|
|
from pathlib import Path
|
|
|
|
import numpy as np
|
|
import PIL.Image
|
|
import torch
|
|
|
|
logger = logging.getLogger(__name__)
|
|
|
|
|
|
def safe_stop_image_writer(func):
|
|
def wrapper(*args, **kwargs):
|
|
try:
|
|
return func(*args, **kwargs)
|
|
except BaseException:
|
|
dataset = kwargs.get("dataset")
|
|
writer = getattr(dataset, "writer", None) if dataset else None
|
|
if writer is not None and writer.image_writer is not None:
|
|
logger.warning("Waiting for image writer to terminate...")
|
|
writer.image_writer.stop()
|
|
raise
|
|
|
|
return wrapper
|
|
|
|
|
|
def squeeze_single_channel(array: np.ndarray) -> np.ndarray:
|
|
"""Drop a leading or trailing singleton channel dim: ``(1, H, W)`` / ``(H, W, 1)`` -> ``(H, W)``.
|
|
|
|
Unlike ``array.squeeze()``, this only removes the channel axis, never an ``H`` or ``W`` of size 1.
|
|
"""
|
|
if array.ndim == 3:
|
|
if array.shape[0] == 1:
|
|
return array[0]
|
|
if array.shape[-1] == 1:
|
|
return array[..., 0]
|
|
return array
|
|
|
|
|
|
def image_array_to_pil_image(image_array: np.ndarray, range_check: bool = True) -> PIL.Image.Image:
|
|
"""Convert a NumPy array to a PIL Image, preserving precision for grayscale.
|
|
|
|
Behaviour by shape:
|
|
|
|
- ``(H, W)`` or ``(1, H, W)`` / ``(H, W, 1)``: single-channel grayscale.
|
|
The native dtype is preserved using the matching PIL mode
|
|
(``I;16`` / ``F``). This is the path used for raw depth maps (no rescaling, clamping, or downcasting)
|
|
- ``(3, H, W)`` / ``(H, W, 3)``: RGB. Channels-first inputs are transposed
|
|
to channels-last. Float inputs in ``[0, 1]`` are scaled to ``uint8``
|
|
(existing behaviour, gated by ``range_check``).
|
|
|
|
Other shapes / channel counts raise ``NotImplementedError`` or
|
|
``ValueError``.
|
|
"""
|
|
# TODO(CarolinePascal): 4 dimensions RGB-D images
|
|
if image_array.ndim not in (2, 3):
|
|
raise ValueError(f"The array has {image_array.ndim} dimensions, but 2 or 3 is expected for an image.")
|
|
|
|
# Squeeze 3D single-channel inputs to 2D so depth maps work whether the
|
|
# caller emits (H, W), (1, H, W), or (H, W, 1).
|
|
image_array = squeeze_single_channel(image_array)
|
|
|
|
if image_array.ndim == 2:
|
|
if image_array.dtype not in [np.uint16, np.float32]:
|
|
raise ValueError(
|
|
f"Unsupported single-channel image dtype: {image_array.dtype}. "
|
|
f"Supported dtypes: {sorted(str(d) for d in [np.uint16, np.float32])}."
|
|
)
|
|
return PIL.Image.fromarray(np.ascontiguousarray(image_array))
|
|
|
|
# 3D path: must be RGB (3 channels), channels-first or channels-last.
|
|
if image_array.shape[0] == 3:
|
|
# Transpose from pytorch convention (C, H, W) to (H, W, C)
|
|
image_array = image_array.transpose(1, 2, 0)
|
|
|
|
elif image_array.shape[-1] != 3:
|
|
raise NotImplementedError(
|
|
f"The image has {image_array.shape[-1]} channels, but 3 is required for now."
|
|
)
|
|
|
|
if image_array.dtype != np.uint8:
|
|
if range_check:
|
|
max_ = image_array.max().item()
|
|
min_ = image_array.min().item()
|
|
if max_ > 1.0 or min_ < 0.0:
|
|
raise ValueError(
|
|
"The image data type is float, which requires values in the range [0.0, 1.0]. "
|
|
f"However, the provided range is [{min_}, {max_}]. Please adjust the range or "
|
|
"provide a uint8 image with values in the range [0, 255]."
|
|
)
|
|
|
|
image_array = (image_array * 255).astype(np.uint8)
|
|
|
|
return PIL.Image.fromarray(image_array)
|
|
|
|
|
|
def save_kwargs_for_path(fpath: Path, compress_level: int) -> dict:
|
|
"""Pick the right format-specific kwargs for :meth:`PIL.Image.Image.save`.
|
|
|
|
PNG uses ``compress_level`` (0-9, zlib). TIFF uses ``compression`` (raw) for lossless raw depth maps.
|
|
"""
|
|
suffix = Path(fpath).suffix.lower()
|
|
if suffix == ".png":
|
|
return {"compress_level": compress_level}
|
|
if suffix in (".tif", ".tiff"):
|
|
return {"compression": "raw"}
|
|
else:
|
|
raise ValueError(f"Unsupported image file extension: {suffix}")
|
|
|
|
|
|
def write_image(image: np.ndarray | PIL.Image.Image, fpath: Path, compress_level: int = 1):
|
|
"""
|
|
Saves a NumPy array or PIL Image to a file.
|
|
|
|
This function handles both NumPy arrays and PIL Image objects, converting
|
|
the former to a PIL Image before saving. It includes error handling for
|
|
the save operation. The output format is inferred from the *fpath*
|
|
extension: ``.png`` → PNG with ``compress_level``, ``.tiff`` / ``.tif``
|
|
→ lossless raw depth maps (TIFF).
|
|
|
|
Args:
|
|
image (np.ndarray | PIL.Image.Image): The image data to save.
|
|
fpath (Path): The destination file path for the image.
|
|
compress_level (int, optional): The compression level for the saved
|
|
image, as used by PIL.Image.save(). Defaults to 1.
|
|
Refer to: https://github.com/huggingface/lerobot/pull/2135
|
|
for more details on the default value rationale.
|
|
|
|
Raises:
|
|
TypeError: If the input 'image' is not a NumPy array or a
|
|
PIL.Image.Image object.
|
|
|
|
Side Effects:
|
|
Logs an error message if the image writing process fails for any reason.
|
|
"""
|
|
try:
|
|
if isinstance(image, np.ndarray):
|
|
img = image_array_to_pil_image(image)
|
|
elif isinstance(image, PIL.Image.Image):
|
|
img = image
|
|
else:
|
|
raise TypeError(f"Unsupported image type: {type(image)}")
|
|
img.save(fpath, **save_kwargs_for_path(fpath, compress_level))
|
|
except Exception as e:
|
|
logger.error("Error writing image %s: %s", fpath, e)
|
|
|
|
|
|
def worker_thread_loop(queue: queue.Queue):
|
|
while True:
|
|
item = queue.get()
|
|
if item is None:
|
|
queue.task_done()
|
|
break
|
|
image_array, fpath, compress_level = item
|
|
write_image(image_array, fpath, compress_level)
|
|
queue.task_done()
|
|
|
|
|
|
def worker_process(queue: queue.Queue, num_threads: int):
|
|
threads = []
|
|
for _ in range(num_threads):
|
|
t = threading.Thread(target=worker_thread_loop, args=(queue,))
|
|
t.daemon = True
|
|
t.start()
|
|
threads.append(t)
|
|
for t in threads:
|
|
t.join()
|
|
|
|
|
|
class AsyncImageWriter:
|
|
"""
|
|
This class abstract away the initialisation of processes or/and threads to
|
|
save images on disk asynchronously, which is critical to control a robot and record data
|
|
at a high frame rate.
|
|
|
|
When `num_processes=0`, it creates a threads pool of size `num_threads`.
|
|
When `num_processes>0`, it creates processes pool of size `num_processes`, where each subprocess starts
|
|
their own threads pool of size `num_threads`.
|
|
|
|
The optimal number of processes and threads depends on your computer capabilities.
|
|
We advise to use 4 threads per camera with 0 processes. If the fps is not stable, try to increase or lower
|
|
the number of threads. If it is still not stable, try to use 1 subprocess, or more.
|
|
"""
|
|
|
|
def __init__(self, num_processes: int = 0, num_threads: int = 1):
|
|
self.num_processes = num_processes
|
|
self.num_threads = num_threads
|
|
self.queue = None
|
|
self.threads = []
|
|
self.processes = []
|
|
self._stopped = False
|
|
|
|
if num_threads <= 0 and num_processes <= 0:
|
|
raise ValueError("Number of threads and processes must be greater than zero.")
|
|
|
|
if self.num_processes == 0:
|
|
# Use threading
|
|
self.queue = queue.Queue()
|
|
for _ in range(self.num_threads):
|
|
t = threading.Thread(target=worker_thread_loop, args=(self.queue,))
|
|
t.daemon = True
|
|
t.start()
|
|
self.threads.append(t)
|
|
else:
|
|
# Use multiprocessing
|
|
self.queue = multiprocessing.JoinableQueue()
|
|
for _ in range(self.num_processes):
|
|
p = multiprocessing.Process(target=worker_process, args=(self.queue, self.num_threads))
|
|
p.daemon = True
|
|
p.start()
|
|
self.processes.append(p)
|
|
|
|
def save_image(
|
|
self, image: torch.Tensor | np.ndarray | PIL.Image.Image, fpath: Path, compress_level: int = 1
|
|
):
|
|
if isinstance(image, torch.Tensor):
|
|
# Convert tensor to numpy array to minimize main process time
|
|
image = image.cpu().numpy()
|
|
self.queue.put((image, fpath, compress_level))
|
|
|
|
def wait_until_done(self):
|
|
self.queue.join()
|
|
|
|
def stop(self):
|
|
if self._stopped:
|
|
return
|
|
|
|
if self.num_processes == 0:
|
|
for _ in self.threads:
|
|
self.queue.put(None)
|
|
for t in self.threads:
|
|
t.join()
|
|
else:
|
|
num_nones = self.num_processes * self.num_threads
|
|
for _ in range(num_nones):
|
|
self.queue.put(None)
|
|
for p in self.processes:
|
|
p.join()
|
|
if p.is_alive():
|
|
p.terminate()
|
|
self.queue.close()
|
|
self.queue.join_thread()
|
|
|
|
self._stopped = True
|