From MDPs and the Bellman equation to model-free control: how value-based methods learn Q-functions through Monte Carlo, TD, SARSA, and Q-learning, and how policy-gradient methods like REINFORCE optimize the policy directly — with a variance-reducing baseline.
A gentle introduction to Reinforcement Learning
