Jonathan's Interactive Research Labs

Interactive Demos & Algorithm Playgrounds

Hands-on web simulations, playable RL agents, and visual decision engines built from my research papers and software repositories.

Reinforcement Learning

Permuted MNIST: Loss of Plasticity vs. SWR & CBP

Live per-task online accuracy playground demonstrating Loss of Plasticity across 50 sequential tasks, comparing Vanilla SGD against SWR and Continual Backprop (CBP).

Loss of PlasticitySWRContinual Backpropmodular-rl
Reinforcement Learning

Plasticity Correlates Dashboard: Dead Units, Rank & Gradients

Diagnostic 4-chart dashboard tracking physical neural correlates over tasks: Dead Units %, Weight Magnitude, Gradient Norm, and Representation Stable Rank.

Dead UnitsStable RankVanishing Gradientsmodular-rl
Reinforcement Learning

Baird's Counterexample: TD Divergence vs. GTD/TDC

Interactive split-screen demo of the Deadly Triad of RL. Watch semi-gradient TD(0) diverge exponentially to ∞ while Gradient TD (GTD0 / TDC) remains stable.

GTD(0)TDC / Fast-GTDOff-Policymodular-rl
Reinforcement Learning

TD(λ) Eligibility Trace Explorer & Scrubber

Educational single-step scrubber revealing how eligibility trace vectors accumulate, decay, and assign temporal credit backward across states.

TD(λ)Eligibility TracesCredit Assignmentmodular-rl
Reinforcement Learning

Random Walk Value Estimation & TD(λ)

Animate 5-state Random Walk value estimation from Sutton (1988). Interactive sliders for α and λ, TD(0) vs TD(λ) vs Monte Carlo toggles, and live bias vs. variance tradeoff analysis.

TD(λ)Temporal DifferenceMonte Carlomodular-rl
Math & Monte Carlo

Mississippi Marbles Strategy & Monte Carlo Simulator

Play Mississippi Marbles manually or toggle the optimal dynamic programming / Monte Carlo expected value (EV) policy solver to visualize optimal decision boundaries.

Monte CarloExpected ValueProbabilityReact Canvas