Codex deep diagnosis surfaced two real issues in our MCTS pipeline:
1. mcts.pyx::search() hardcoded prepare_simulation_batch(state, traverser, 1)
instead of respecting MctsConfig.parallel_simulations. Standalone evals
(eval_checkpoint, evaluate_with_mcts sequential path, eval_worker) all
use this entry point, so all eval-time MCTS was running 1 sim per batch
regardless of the configured 64. Training was unaffected because it
goes through interleaved_self_play._run_search_jobs which respects the
config. Fix uses min(config.parallel_simulations, sims - completed).
2. use_rollout_value defaulted to True (config.py) but was never set in
the YAML. With this, _expand_with_prior returns the heuristic rollout
value and discards network_value, so the network value head is trained
from final game scores but its outputs are never fed back into MCTS
backups. This explains why mcts/value_prediction_error stays high
despite training -- learning the value head produces no behavioral
change because MCTS never reads it.
Now setting use_rollout_value=false in default.yaml so the network value
head closes the loop. Combined with the existing Dirichlet root noise +
heuristic rollout removal, this should give the network's value learning
actual leverage on action selection.
Also: updated test_search_visit_counts_match_with_parallel_simulations
to test the correct invariant (legal-action set match + total visit
count near n_sims) rather than literal visit-count equality, which was
only true under the previous bug.
Tests: 19/19 passing.
Key changes for ISMCTS speed and correctness:
- Cython port: HeuristicBot helpers (`heuristic_cy.pyx` + new `.pxd`) and
ISMCTS searcher (`mcts.pyx`) now run as cdef. Both share a fast
unified-action path through GameState's C interface to avoid Python
round-trips on hot rollout/tree-walk paths.
- Multi-process self-play and eval: `workers.py`, `eval_worker.py`,
`interleaved_self_play.py`, plus trainer wiring with ProcessPoolExecutor
+ spawn context. Eval inside `evaluate.py` is parallel per opponent.
- ISMCTS-specific eval (`evaluate.py`) runs MCTS at decision time so the
metric matches deploy mode; `evaluation.eval_with_mcts` flag preserves
backwards-compatible policy-only eval when needed.
- Trainer logs progress per phase (self-play start/done, eval per
opponent), and value loss is now scaled by `value_scale` so policy and
value losses sit on comparable magnitudes.
- Compact info-set key (`info_set.py`) using packed-struct format and
child-key reuse during MCTS descent to cut per-step canonicalization.
Tests: 19 ISMCTS suite passing, including parity (Cython-vs-Python
sequential, batched-vs-sequential visit counts, push/pop round-trip).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Wrap AlphaZeroNet with a logits-only view so IS-MCTS training evaluation can call evaluate_strategy_network and emit the same full diagnostic metric set as Deep CFR. Adds root prior capture and per-iteration MCTS entropy, value error, and policy-vs-search KL metrics.
Tests: uv run python -m pytest tests/games/classic/ismcts/ -x; uv run python -m pytest tests/games/classic/test_deep_cfr_trainer.py -x; uv run lost-cities-ismcts train --config configs/ismcts/mini.yaml --set run.experiment_name=ismcts-metrics-smoke --set run.max_iterations=2 --set training.games_per_iter=2
Implements a proof-of-concept single-observer IS-MCTS trainer with AlphaZero-style policy/value network, determinization, replay, self-play, CLI configs, and focused tests. Mini acceptance run reaches positive random eval while keeping play_action_rate above the Deep CFR trap threshold.
Tests: uv run python -m pytest tests/games/classic/ismcts/ -x; uv run python -m pytest tests/games/classic/test_deep_cfr_trainer.py -x; uv run lost-cities-ismcts train --config configs/ismcts/mini.yaml