Neutreeko

An abstract two-player board game I've started using to experiment with reinforcement learning. You can play against the AI right in your browser.

Prototype · in progress

Work in progress

I started this personal project in my spare time and have only spent a little time on it so far: it's a first draft, still far from a polished version. It's mainly a way for me to learn, especially about neural networks and reinforcement learning. For now, the trained agent loses to the simplest minimax search.

See the next steps

Game · you vs the AI

Prototype

You0

Draws0

AIIntermediate0

Your turn: pick one of your green pawns.

AI level
Who starts?
  • Your pawns
  • AI pawns
  • Last move

Rules

How to play

  1. Slide all the way

    On your turn, move one of your three pawns in any of the eight directions. It slides in a straight line until it hits the edge of the board or another pawn: no stopping halfway.

  2. Line up your three pawns

    The first player to place their three pawns next to each other, in a row, a column or a diagonal, wins.

  3. Draws

    If the same position comes up three times, the game is a draw.

Click one of your green pawns, then a square marked with a dot. With a keyboard: arrow keys to move around the board, Enter to choose, Escape to cancel.

Under the hood

Three opponents, three approaches

The original game is written in Python. For this site, the game logic and the neural network's inference were ported to JavaScript, then checked against the Python version on 120 test positions. Everything runs in your browser: no data is sent anywhere.

Level 1

Beginner

DQN network alone

A neural network (75 → 128 → 64 → 24) trained with Deep Q-Learning. It plays the move it rates highest, without looking ahead. Still learning, it is the easiest opponent.

Level 2

Intermediate

DQN network + safety net

The same network, framed by two hand-written rules: play the winning move when there is one, and rule out moves that hand the opponent a win on the next turn.

Level 3

Expert

Alpha-beta minimax search

No learning at all: a search that looks four plies ahead and scores positions by how close and how aligned the pawns are. It is the benchmark the agent has to learn to beat.

Indicative measurementsCurrent version
Beginner vs 1-ply minimax
0 wins, 37 losses and 3 draws over 40 games
Intermediate vs 1-ply minimax
18 wins, 20 losses and 2 draws over 40 games
Expert vs 2-ply minimax
20 wins over 20 games

Program-versus-program matches, alternating who starts.

Next steps

A few ideas I'd like to try as I keep learning:

  • training it against the minimax search at increasing depths, rather than only against itself;
  • using the board's eight symmetries to multiply the training data;
  • pairing the network with a tree search (MCTS, AlphaZero-style) so it can look several moves ahead.

Goal: beat the minimax search without a hand-written safety net.

Source code

The code is on GitHub

Game engine, DQN training and minimax search: it is all in the repository.