Animate tabular Temporal Difference learning on the 5-state Random Walk MDP. Explore convergence of value estimates V(s) → V*(s), adjust learning rate α and trace decay λ, and visualize the live bias-variance tradeoff.
Agent starts at C. Step transitions walk Left/Right with 50% probability.
Estimated state values (colored) converging to analytical true values (striped).
Live error decay curve tracking RMS error over episodes.