Beginner
DQN network alone
A neural network (75 → 128 → 64 → 24) trained with Deep Q-Learning. It plays the move it rates highest, without looking ahead. Still learning, it is the easiest opponent.
An abstract two-player board game I've started using to experiment with reinforcement learning. You can play against the AI right in your browser.
Work in progress
I started this personal project in my spare time and have only spent a little time on it so far: it's a first draft, still far from a polished version. It's mainly a way for me to learn, especially about neural networks and reinforcement learning. For now, the trained agent loses to the simplest minimax search.
See the next stepsYou0
Draws0
AIIntermediate0
Your turn: pick one of your green pawns.
Rules
On your turn, move one of your three pawns in any of the eight directions. It slides in a straight line until it hits the edge of the board or another pawn: no stopping halfway.
The first player to place their three pawns next to each other, in a row, a column or a diagonal, wins.
If the same position comes up three times, the game is a draw.
Click one of your green pawns, then a square marked with a dot. With a keyboard: arrow keys to move around the board, Enter to choose, Escape to cancel.
Under the hood
The original game is written in Python. For this site, the game logic and the neural network's inference were ported to JavaScript, then checked against the Python version on 120 test positions. Everything runs in your browser: no data is sent anywhere.
DQN network alone
A neural network (75 → 128 → 64 → 24) trained with Deep Q-Learning. It plays the move it rates highest, without looking ahead. Still learning, it is the easiest opponent.
DQN network + safety net
The same network, framed by two hand-written rules: play the winning move when there is one, and rule out moves that hand the opponent a win on the next turn.
Alpha-beta minimax search
No learning at all: a search that looks four plies ahead and scores positions by how close and how aligned the pawns are. It is the benchmark the agent has to learn to beat.
Program-versus-program matches, alternating who starts.
A few ideas I'd like to try as I keep learning:
Goal: beat the minimax search without a hand-written safety net.
Source code
Game engine, DQN training and minimax search: it is all in the repository.