<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Next-Token Prediction on /home/vigi99</title><link>https://viig99.github.io/tags/next-token-prediction/</link><description>Recent content in Next-Token Prediction on /home/vigi99</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Thu, 16 Jul 2026 17:15:10 -0400</lastBuildDate><atom:link href="https://viig99.github.io/tags/next-token-prediction/index.xml" rel="self" type="application/rss+xml"/><item><title>Building a Chess Bot with HSTU: From Lichess Pretraining to Value Search</title><link>https://viig99.github.io/docs/posts/imba-chess/</link><pubDate>Thu, 16 Jul 2026 00:00:00 +0000</pubDate><guid>https://viig99.github.io/docs/posts/imba-chess/</guid><description>&lt;h2 id="build-log-snapshot"&gt;Build Log Snapshot&lt;a class="anchor" href="#build-log-snapshot"&gt;#&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;This is my running build log for &lt;a href="https://github.com/viig99/imba-chess" target="_blank" rel="noopener" &gt;&lt;code&gt;imba-chess&lt;/code&gt;&lt;/a&gt;. I update it as things change, so it reads out of order in places. That is on purpose.&lt;/p&gt;
&lt;p&gt;Where it stands today:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;I trained a small HSTU sequence model to imitate high-Elo Lichess games (next-move prediction). It reaches &lt;code&gt;hr@10&lt;/code&gt; around &lt;code&gt;0.92&lt;/code&gt; and &lt;code&gt;top-1&lt;/code&gt; around &lt;code&gt;0.43&lt;/code&gt; on held-out games.&lt;/li&gt;
&lt;li&gt;On its own, that policy is weak at actual chess. Greedy play scores &lt;code&gt;0.21&lt;/code&gt; against Stockfish limited to 1400 Elo. Good imitation, bad chess.&lt;/li&gt;
&lt;li&gt;Adding a WDL value head and a value-guided search at move-selection time changes the story. The current best system scores &lt;code&gt;0.595&lt;/code&gt; against Stockfish at 2200 Elo, which puts the whole thing at roughly &lt;strong&gt;2250 Elo&lt;/strong&gt;. That is about an 800 to 1000 Elo jump over the raw policy, with no reinforcement learning and no self-play yet.&lt;/li&gt;
&lt;li&gt;The search that got me there is &lt;code&gt;value_search_halving&lt;/code&gt;: sequential halving to decide which root move to spend compute on, plus a prior-ordered beam for the tree underneath it.&lt;/li&gt;
&lt;li&gt;The current bet is distilling that search back into the policy head, so the policy itself gets stronger without paying for search at inference. This is grounded in the Gumbel MuZero policy-improvement result. It is designed and being built, not yet a proven win.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Two model sizes show up below: &lt;code&gt;v3&lt;/code&gt; (512-dim, 6 layers, about 10M params) and &lt;code&gt;v4&lt;/code&gt; (768-dim, 8 layers, about 27M params). The bigger &lt;code&gt;v4&lt;/code&gt; trunk finally hit my original 20-30M target and it mattered. Most of the dev loop runs on modest hardware: an 8GB laptop GPU (RTX 3070 class), some runs on a 4090, and a rented 5090 for a while to speed up rollout generation.&lt;/p&gt;</description></item></channel></rss>