50-iter sweep on default.yaml with stronger MCTS exploration:
- c_puct: 3.0 -> 5.0 (UCB weight, more exploration of low-prior actions)
- root_dirichlet_epsilon: 0.25 -> 0.4 (more noise injected at root prior)
Standalone eval at iter 50 (30 games/opponent, all natural-end, timeouts=0):
- vs heuristic-balanced: W=0/30 S=-70.5 (PA 0.15)
- vs heuristic-aggressive: W=2/30 S=-65.1 (PA 0.14) [+10, +4]
- vs heuristic-cautious: W=1/30 S=-48.0 (PA 0.14) [+2]
3 natural-end wins vs prev trapfix baseline iter 44 (which had 0 natural
wins + 1 timeout-tie). Stall trap fixed remains true (timeouts=0 in c1).
Trade-off observed: more exploration -> higher variance. Score avg vs
cautious worsened (-32 -> -48), but win events appeared. For the
non-terminal-win objective, exploration win > score-avg loss.
Next: commit to long run (300 iter) with these params before tuning more.