Policies trained from data instead of hand-written controllers.
Robust quadruped locomotion via multiplicity of behavior.