Datasets that feed robot vision and SLAM.
3,600 hours of egocentric human video used to pretrain manipulation models.