Policies learned through trial and reward.
RL for humanoid locomotion with zero-shot sim-to-real transfer.
Robust quadruped locomotion via multiplicity of behavior.