Building a Chess Bot with HSTU: From Lichess Pretraining to Value Search

Build Log Snapshot#

This is my running build log for imba-chess. I update it as things change, so it reads out of order in places. That is on purpose.

Where it stands today:

  • I trained a small HSTU sequence model to imitate high-Elo Lichess games (next-move prediction). It reaches hr@10 around 0.92 and top-1 around 0.43 on held-out games.
  • On its own, that policy is weak at actual chess. Greedy play scores 0.21 against Stockfish limited to 1400 Elo. Good imitation, bad chess.
  • Adding a WDL value head and a value-guided search at move-selection time changes the story. The current best system scores 0.595 against Stockfish at 2200 Elo, which puts the whole thing at roughly 2250 Elo. That is about an 800 to 1000 Elo jump over the raw policy, with no reinforcement learning and no self-play yet.
  • The search that got me there is value_search_halving: sequential halving to decide which root move to spend compute on, plus a prior-ordered beam for the tree underneath it.
  • The current bet is distilling that search back into the policy head, so the policy itself gets stronger without paying for search at inference. This is grounded in the Gumbel MuZero policy-improvement result. It is designed and being built, not yet a proven win.

Two model sizes show up below: v3 (512-dim, 6 layers, about 10M params) and v4 (768-dim, 8 layers, about 27M params). The bigger v4 trunk finally hit my original 20-30M target and it mattered. Most of the dev loop runs on modest hardware: an 8GB laptop GPU (RTX 3070 class), some runs on a 4090, and a rented 5090 for a while to speed up rollout generation.