Skip to content

Latest commit

 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Approximate value iteration in selfplay

Python 3.12+ JAX Ruff

This repository contains the minimal training and evaluation implementation for the paper, together with the configuration files, benchmark data, checkpoints and solver packages.

Installation

Install the core training and evaluation package with:

pip install .

Optional capabilities are available as extras: accelerated (Triton GPU kernels) and test. The test extra dependencies required to run the tests.

JAX accelerator wheels must be installed explicitly. For an NVIDIA GPU with CUDA 13, install the core package and CUDA-enabled JAX with:

pip install .
pip install --upgrade 'jax[cuda13]'

AlphaZero can use the pure-JAX search implementation, but the Triton kernels are substantially faster on NVIDIA GPUs:

pip install '.[accelerated]'

The C4 and Hex oracle bindings are separate native packages; install them from their respective directories in solvers/ when oracle evaluation is needed.

Training

Training is launched with an algorithm, configuration, and experiment name:

JAX_ENABLE_X64=1 python train.py avi c4 my-experiment
python train.py az hex7x7 my-experiment

(Connect 4 uses 64-bit bitboards and requiresJAX_ENABLE_X64=1, leave x64 disabled for Hex, Othello, and Go.)

Configuration overrides use OmegaConf dot-list syntax:

python train.py avi c4 my-experiment train.lr=1e-4 global.seed=1

Checkpoints are written under ckpts/<algorithm>/<experiment>/<run>/.

Evaluation

Evaluate a saved Equinox checkpoint independently of training:

python evaluate.py CONFIG CHECKPOINT_PATH --mode MODE

Evaluation configs live under config/eval/; pass the filename without its .toml suffix. Released-checkpoint configs are provided for every experiment. The supported modes are dataset, oracle and head-to-head.

Examples:

JAX_ENABLE_X64=1 python evaluate.py c4_avi models/connect_four_avi.eqx \
  --mode dataset

python evaluate.py hex7x7_avi models/hex_7x7_avi.eqx \
  --mode dataset

JAX_ENABLE_X64=1 python evaluate.py c4_avi models/connect_four_avi.eqx \
  --mode head-to-head \
  --opponent-checkpoint models/connect_four_az_n512.eqx

python evaluate.py othello_avi_cross models/othello_avi.eqx \
  --mode head-to-head \
  --opponent-checkpoint models/othello_az_n200.eqx

Oracle mode additionally requires the corresponding native solver package. Hex oracle evaluation can be very slow; fixed-dataset and head-to-head modes are better smoke tests.

Checkpoints

Pre-trained Equinox checkpoints are available from the project website.

Game Algorithm Checkpoint SHA256
Connect Four AlphaZero, seed 0 (512 simulations) connect_four_az_n512.eqx_.zip dc58956901cdfa8e91c713cb910a6f756291f0c6679c550afed104a746aa02da
Connect Four AVI, seed 0 connect_four_avi.eqx_.zip 102eaa194009b59dd40bb7541261c2ecf9798cde936d23301f0a57f25bb20a4b
Hex 7×7 AlphaZero, seed 0 (512 simulations) hex_7x7_az_n512.eqx_.zip ad454377a05733609cd1fce6629d9d22b2fbe4529996864c7cdd70adcd62e7d2
Hex 7×7 AVI, seed 0 hex_7x7_avi.eqx_.zip ec3a1b242fd7ff59e56ebbdcf139501ac66d3222b1d85739acda806941ebf779
Othello AVI, seed 0 othello_avi.eqx_.zip 82f9f887967fad5f9845767ea8482c6b0ce89451fb88ceda8e63c6821a61dac4
Othello MiniZero AZ (200 simulations) othello_az_n200.eqx_.zip b7acc2f3d50954c94dc82ede6b189d35e25d06e312fc5a9a630ceefa74a790ba
Go 9×9 AVI, seed 0 go_9x9_avi.eqx_.zip 506bf1079b16243c102730ebefeda5168e242958ba08f18fa8a7803deb4da7c3
Go 9×9 MiniZero AZ (200 simulations) go_9x9_az_n200.eqx_.zip 596f46e49482eac19d7c00f5473665009f2ee6bf81bc80bedc4a571cf6f1ff93

With the released checkpoints and seed 0, useful sanity checks are C4 random MAE around 0.27 for AVI versus 0.49 for AZ-512, and Hex MAE around 0.23 for AVI versus 0.47 for AZ-512. Head-to-head scores are seed-specific: the paper reports aggregates across paired runs, so a single released pair need not equal the reported mean.

Datasets

The repository includes the benchmark data under data/. Go and Othello opening positions are generated deterministically at evaluation time.

Solvers (oracle)

Solvers sources are located under solvers/. The C4 and Hex solvers are implemented in C++ and wrapped with Python bindings. Requirements and installation instructions are provided in the corresponding README.md files.

Minizero

Go(9x9) and Othello support comparison with MiniZero, as no exact oracles are available for these games. To validate the translation itself, install the test requirements and download the original PyTorch checkpoints:

pip install torch==2.12.1 --index-url https://download.pytorch.org/whl/cpu
pip install '.[test]'
wget -O- https://rlg.iis.sinica.edu.tw/papers/minizero/assets/models/go_9x9.tar.gz \
  | tar -xz -C models/minizero go_9x9_az_n200.pt
wget -O- https://rlg.iis.sinica.edu.tw/papers/minizero/assets/models/othello_8x8.tar.gz \
  | tar -xz -C models/minizero othello_az_n200.pt
python -m pytest tests/test_minizero_translation.py

Citation

If you use this work please cite:

@misc{boige2026surprisingeffectivenessapproximatevalue,
      title={The Surprising Effectiveness of Approximate Value Iteration in Self-Play}, 
      author={Raphael Boige and Amine Boumaza and Bruno Scherrer},
      year={2026},
      eprint={2609.09094},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2609.09094}, 
}

About

Train self-play agents for games using Approximate Value Iteration

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages