Current work
What I am working on now
I am currently working across three connected areas: planning and reinforcement learning for energy control, reliable reinforcement-learning experiments, and reproducible software for multi-agent LLM research. The common thread is making sequential AI systems easier to evaluate, audit and explain.
In progress
Active workstreams
Workstream 1
Planning for energy-aware control
Research question
How can a controller look ahead, meet time and comfort constraints, and avoid unnecessary energy use when future conditions are uncertain?
Current focus
Comparing planning and reinforcement-learning approaches for building heating and domestic hot-water control, using MCTS, PPO, SAC and BOPTEST.
Why it matters
A useful controller must balance efficiency with operational reliability; average reward alone does not show whether deadlines or constraints were actually met.
Workstream 2
Reliable reinforcement learning
Research question
Which numerical and experimental details can change policy optimisation, and how can those effects be isolated without overstating the evidence?
Current focus
Running controlled comparisons, checking policy and task behaviour separately, and keeping a traceable route from each conclusion to its supporting run.
Why it matters
Small implementation details can undermine a comparison if they are not measured and reported explicitly.
Public scope: Only the broad public research direction is described here. Details tied to anonymous review or unpublished results are withheld until public.
Workstream 3
Reproducible AI research software
Research question
How can multi-agent LLM experiments be configured, logged and repeated so that management-research claims remain inspectable?
Current focus
Building experiment software with explicit configuration, structured logging and auditable run outputs for multi-agent LLM studies.
Why it matters
Reproducible infrastructure makes it easier for collaborators to compare runs, find errors and understand which evidence supports a conclusion.
Research practice
How I evaluate the work
Define the decision
State the task, constraints, baselines and evidence needed before running experiments.
Control the comparison
Change one intended factor at a time and keep evaluation conditions aligned.
Audit the run
Record configuration, outputs and diagnostics that separate numerical survival, policy health and task performance.
Communicate the limit
Explain what was measured, what is inferred and what remains unknown.
Public scope
What this page deliberately leaves out
This is a public progress page, not a lab notebook. It gives enough context to explain the direction of the work without exposing private, employer-confidential or anonymous-review material.
- No employer-confidential data, private collaborator details or unpublished outcomes.
- No anonymous-review identity, venue or result details before they are public.
- Dates and role descriptions follow the verified public profile and CV.
Follow the public work
For stable records, use the publications, Work page and verified CV. For a question about the current direction, email me.