VoxPoser
LLM-written 3D value maps for zero-shot manipulation. · By Stanford
VoxPoser uses a language model to write code that composes vision-model outputs into 3D affordance and constraint maps, then plans trajectories through them, performing tasks it
Description
VoxPoser uses a language model to write code that composes vision-model outputs into 3D affordance and constraint maps, then plans trajectories through them, performing tasks it was never trained on.