Policies trained from data instead of hand-written controllers.
RL for humanoid locomotion with zero-shot sim-to-real transfer.
Robust quadruped locomotion via multiplicity of behavior.