Distributions
A Distribution is a distribution over the parameter vector of a policy. It is the object the black-box
optimization algorithms learn: at the beginning of an episode they sample a parameter vector from it, run the
episode with the resulting policy, and update the distribution from the return that was obtained.
A distribution can be contextual, i.e. conditioned on a context vector built from the initial state and the episode info, in which case sampling and updating are performed per context.
Interface for Distributions to represent a generic probability distribution. |
|
Gaussian distribution with fixed covariance matrix. |
|
Gaussian distribution with diagonal covariance matrix. |
|
Gaussian distribution with full covariance matrix. |