Train My Robot
Watch a robot learn to navigate a maze through trial, error, rewards, and penalties — just like real AI agents learn!
How Reinforcement Learning Works
Environment
The maze is the world the robot lives in — walls, open paths, a goal, and traps!
Actions & Rewards
The robot tries moves. Reaching the goal = +10 points. Hitting walls or traps = penalty.
Q-Learning
Robot builds a "Q-Table" — a cheat sheet of which action is best from every position.
Getting Smarter
Over hundreds of episodes, the robot improves — fewer steps, more rewards each run!
Choose Your Maze
📋 Maze Info
Manual Navigation
🎮 Your Score
Robot Training
🤖 Training Stats
📈 Learning Curve — Reward per Episode
🗂️ Q-Table Heatmap (hover cells to inspect best action)
Trained Robot — Best Policy
📊 Full Learning Curve
RL Scientist Badge Unlocked!
You trained a Q-Learning robot to navigate a maze. Enter your name to get your certificate!
Optional. Stays on this device only — not sent to WhizzStep.
Key Concepts You've Mastered
🤖 The Robot
The learner that takes actions in the environment to maximise reward over time.
🗺️ The Maze
The world the agent lives in. It responds to actions with new states and rewards.
📋 The Cheat Sheet
A table that stores the expected reward for every (state, action) pair. Updated after every step.
🎲 Try New vs Use Best
Epsilon (ε) controls this. High ε = explore randomly. Low ε = use learned knowledge.
🧮 The Update Rule
Q(s,a) ← Q(s,a) + α[r + γ·max Q(s',a') − Q(s,a)]. The heart of Q-learning!
🗺️ The Best Path
After training, the optimal policy selects the best action from every state — the arrows you saw!
About this lab
Learning objective: Guide a robot through a maze and observe how rewards and penalties shape behaviour in a simplified Q-learning demonstration.
What this simplifies: A small deterministic maze and a simplified reward model are used; this is not a general-purpose robot-learning system.
Privacy: No learner input leaves the device.
Teacher prompt: Ask the class why this simulation might mislead someone who takes it too literally.
Reflect: What is one thing this activity showed you that you did not expect?
← Back to all Labs