OpenVLA
An open-source 7B vision-language-action model for robot manipulation. · By Stanford
OpenVLA is trained on 970k robot episodes from Open X-Embodiment. It maps images and language instructions to robot actions and can be fine-tuned on consumer GPUs for new robots
Description
OpenVLA is trained on 970k robot episodes from Open X-Embodiment. It maps images and language instructions to robot actions and can be fine-tuned on consumer GPUs for new robots and tasks.