Turning sensor data into objects, poses and maps.
Segment Anything in images and videos.
Monocular depth estimation foundation model.
Deploy vision models (YOLO, SAM, CLIP) on edge devices for robots.