Turning sensor data into objects, poses and maps.
Segment Anything in images and videos.
Monocular depth estimation foundation model.
Deploy vision models (YOLO, SAM, CLIP) on edge devices for robots.
ROS driver for ReSpeaker microphone arrays.