Back to Research Labs
modular-rl
Sutton (1988) Library Namesake Demo

Random Walk Value Estimation & TD(λ)

Animate tabular Temporal Difference learning on the 5-state Random Walk MDP. Explore convergence of value estimates V(s) → V*(s), adjust learning rate α and trace decay λ, and visualize the live bias-variance tradeoff.

Algorithm Paradigm
TDLAMBDA

Hyperparameters

Learning Rate (α):0.100
0.005 (Slow)0.250.50 (Aggressive)
Trace Decay (λ):0.50
0.0 (TD(0))0.5 (Balanced)1.0 (Monte Carlo)
Step Delay (ms):50 ms

Execution Controls

5-State Random Walk MDP Pathway

Agent starts at C. Step transitions walk Left/Right with 50% probability.

Episodes: 0
Steps: 0
0.0
Terminal L
A0.50
e: 0.00
B0.50
e: 0.00
C0.50
e: 0.00
D0.50
e: 0.00
E0.50
e: 0.00
1.0
Terminal R
Current State: C
True Values: [0.167, 0.333, 0.500, 0.667, 0.833]

Value Estimates V(s) vs. True Values V*(s)

Estimated state values (colored) converging to analytical true values (striped).

Current RMS Error:0.2357
State A
Est: 0.500True: 0.167Δ: 0.333
State B
Est: 0.500True: 0.333Δ: 0.167
State C
Est: 0.500True: 0.500Δ: 0.000
State D
Est: 0.500True: 0.667Δ: 0.167
State E
Est: 0.500True: 0.833Δ: 0.333

RMS Error Trajectory Over Episodes

Live error decay curve tracking RMS error over episodes.