Repository for Probabilistic AI, featuring implementations of Gaussian Process Regression, Bayesian Neural Networks, Bayesian Optimization, and Actor-Critic Reinforcement Learning.
The task was to implement an algorithm that, by practicing on a simulator, learns a control policy for the Lunar Lander problem. The method suggested is a variant of policy gradient with two additional features, namely (1) Rewards-to-go, and (2) Generalized Advantage Estimatation, both aiming at decreasing the variance of the policy gradient estimates while keeping them unbiased.