Policies learned through trial and reward.
Reliable PyTorch implementations of RL algorithms.
Single-file, research-friendly deep RL implementations.
Robust quadruped locomotion via multiplicity of behavior.