loading…
A4.3.6 HL · reinforcement learning
The agent learns from rewards as it moves through the grid. Reaching the goal gives a reward, entering the pit gives a penalty, and other moves have a small cost. Shorter successful routes give more reward. Change how often the agent explores and watch its learned route develop.
The open circle marks the start. Green marks the goal and orange marks the pit. Stronger shading shows a higher learned estimate of future reward, with later rewards discounted. Arrows show the preferred moves. Click an empty square, or focus it and press Enter, to add a wall. Training restarts for the changed maze.