Turning sensor data into objects, poses and maps.
Segment Anything in images and videos.
Monocular depth estimation foundation model.