Policies learned through trial and reward.
Reliable PyTorch implementations of RL algorithms.
Single-file, research-friendly deep RL implementations.