Pong using Policy Gradients
A policy gradient agent trained from raw pixels to learn how to play Pong through self-play, with no hand-engineered features — just the difference between consecutive frames as input.
Objective
The policy is a neural network parameterized by , trained to maximize the expected discounted return
The REINFORCE gradient
Directly optimizing without a value function uses the REINFORCE estimator:
where is the return from timestep onward. In practice, is centered and normalized across each batch of episodes to reduce the variance of the gradient estimate before the policy update
Result
After enough self-play episodes, the agent reliably beats the built-in Pong AI, learning entirely from the / reward at the end of each point with no intermediate shaping.
Source code: github.com/aarondavis-git/Pong-PolicyGradients