Large pretrained vision models.
Segment Anything in images and videos.
Monocular depth estimation foundation model.