Policies trained from data instead of hand-written controllers.
Reliable PyTorch implementations of RL algorithms.
Single-file, research-friendly deep RL implementations.
Robust quadruped locomotion via multiplicity of behavior.