Cycle 3 prep: fix search() ignoring parallel_simulations + flip use_rollout_value=false

Codex deep diagnosis surfaced two real issues in our MCTS pipeline:

1. mcts.pyx::search() hardcoded prepare_simulation_batch(state, traverser, 1)
   instead of respecting MctsConfig.parallel_simulations. Standalone evals
   (eval_checkpoint, evaluate_with_mcts sequential path, eval_worker) all
   use this entry point, so all eval-time MCTS was running 1 sim per batch
   regardless of the configured 64. Training was unaffected because it
   goes through interleaved_self_play._run_search_jobs which respects the
   config. Fix uses min(config.parallel_simulations, sims - completed).

2. use_rollout_value defaulted to True (config.py) but was never set in
   the YAML. With this, _expand_with_prior returns the heuristic rollout
   value and discards network_value, so the network value head is trained
   from final game scores but its outputs are never fed back into MCTS
   backups. This explains why mcts/value_prediction_error stays high
   despite training -- learning the value head produces no behavioral
   change because MCTS never reads it.

Now setting use_rollout_value=false in default.yaml so the network value
head closes the loop. Combined with the existing Dirichlet root noise +
heuristic rollout removal, this should give the network's value learning
actual leverage on action selection.

Also: updated test_search_visit_counts_match_with_parallel_simulations
to test the correct invariant (legal-action set match + total visit
count near n_sims) rather than literal visit-count equality, which was
only true under the previous bug.

Tests: 19/19 passing.
This commit is contained in:
2026-05-11 06:42:20 +09:00
parent be0c1a8d62
commit 8b7ed66ffd
3 changed files with 24 additions and 5 deletions
@@ -338,12 +338,16 @@ cdef class IsMctsSearcher:
)
cdef int sims = int(n_sims or self.config.n_simulations)
cdef int completed = 0
cdef int batch_size
cdef list pending
cdef list legal
cdef int action
cdef dict result
while completed < sims:
pending = self.prepare_simulation_batch(state, traverser, 1)
batch_size = min(int(self.config.parallel_simulations), sims - completed)
if batch_size <= 0:
batch_size = 1
pending = self.prepare_simulation_batch(state, traverser, batch_size)
if not pending:
break
self.evaluate_and_backup(pending)