Demos
Research is better when you can see it in action. Below are browser-based visualizations of the algorithms I use in my PhD.
Play against a planner
Tic-tac-toe against Monte Carlo Tree Search. Before it moves, the search runs four hundred simulated games from the current position. The number in each empty square is the share of those simulations spent considering that move, and the shaded fill is its estimated win rate — the same selection statistics the hot-water controller uses over heating schedules.
Your move — you are X.
Make a move and the search will run 400 simulated games before replying.
What to look for. The percentages are never spread evenly. Search concentrates on a few candidate moves and barely examines the rest — that uneven allocation is the whole point, and it is what makes tree search affordable on problems where evaluating every option is impossible.
You versus UCB1
Five machines with hidden payout rates. Pull one and UCB1 takes its own turn on the same machines. The bar is each arm's estimated value from your pulls; the figure beneath is how many times UCB1 chose it. Regret is the reward you both gave up by not always playing the best arm.
You have pulled 0 times · UCB1 has pulled 0 times
0
your reward
0
UCB1 reward
0.0
your regret
0.0
UCB1 regret
What to look for. Pull one arm repeatedly and your regret climbs steadily even when that arm is decent — you never learn what you gave up. UCB1 spends early pulls losing on purpose, then reveals the true rates and check whose regret is lower. Every controller I build faces this trade-off: exploiting the schedule that worked yesterday, or testing one that might be cheaper.
Prune the search tree
A three-branch decision tree with fixed leaf rewards. Click any node to cut it out of the search, then watch where the two thousand simulations go instead. Cutting the branch the search likes is how you find out whether it had a second-best answer or only one.
What to look for. Cut the branch marked in red and the visit counts redistribute immediately. Sometimes the second choice is nearly as good; sometimes the estimate falls off a cliff. In a real controller those cuts are not hypothetical — they are the actions a constraint has forbidden, and the gap tells you how much the constraint costs.
Break your own objective
A tank of water must reach sixty degrees by seven o'clock, and electricity is dearest just before then. Set the weights and an exhaustive optimiser searches all 4,096 on-off schedules for the one that maximises your objective. It optimises exactly what you wrote down, which is rarely what you meant.
Bars are water temperature; the strip beneath shows when the heater runs. The deadline column is marked in red.
—
at the deadline
—
electricity cost
—
half-hours heating
—
deadline
Two settings worth trying. Drop the deadline penalty to zero and the tank never reaches temperature — the objective is satisfied by failing. Then push comfort to ten with energy at zero and it boils the water through the costliest hours. Neither is a bug in the optimiser; both are exactly what was asked for, which is why I treat objective specification as part of the experiment rather than a detail.
Draw a world, watch it learn
Click cells to build walls or move the goal, then train a Q-learning agent on the world you drew. Each arrow is the action with the highest value in that cell and the shading is how valuable the cell has become. Change the walls after training to see how much of the learned policy was really about your layout.
Untrained
Clicking a cell adds or removes a wall. The agent starts at S and is rewarded for reaching G.
ε-greedy with decay · α 0.1 · γ 0.95 · step cost 0.04
What to look for. Train until arrows point along a clear route, then wall that route off and train again without forgetting. The old values persist and briefly send the agent into your new wall, because Q-learning has learned about a world that no longer exists. This is why a policy trained on one building model does not transfer to another.
One seed is not a result
Twelve training runs, identical hyperparameters, different random seeds. Each row is one run across sixty evaluation points; darker means better performance. Some runs converge, some plateau, and some collapse outright — the numerical failures that motivate my work on reliable reinforcement learning.
—
best seed
—
median
—
worst seed
—
collapsed
Best seed claims
All seeds claim
What to look for. Rerun several times and watch the two claims diverge. Both are computed from the same twelve runs; only the reporting choice differs. This is the whole argument for reporting spread and failure counts rather than a single curve.
Recover the physics
A tank of water cools and is sometimes heated. You get a hundred noisy measurements and a library of candidate terms; choose which belong in the equation and sparse regression fits the coefficients. The true law uses three terms. Coefficients are fitted on the first fourteen samples only and scored on the remaining eighty-six, so overfitting is visible rather than hidden.
Recovered equation
—
terms used
—
fit error
—
held-out error
What to look for. Select everything and the fit error falls while the held-out error climbs — the extra terms are describing the noise. That gap is the reason I score surrogate models on data they were never fitted to before letting a controller depend on them.