
coolguyandClaude Opus 4.7
bac630b50b
Bump avg-strategy 1000iter config: traversals 4x, LR 3e-5→1e-4, LCFR, lighter eval
Apply consultant + review recommendations to address training-budget
shortfall and underutilized weighting:
- traversal.traversals_per_player: 70 → 280 (4× sample touches/iter to
reduce regret estimate variance early)
- optimization.learning_rate: 3e-5 → 1e-4 (was too low for the 512×1024
updates schedule)
- training_weighting.mode: none → lcfr (faster convergence; alpha/beta/
gamma fields are inert with mode=none)
- evaluation.eval_every: 5 → 25 (eval was costing more wall-clock than
training; 6 opponents × 100 games × 200 evals adds up)
- Drop accidental duplicate keys in traversal/optimization sections
(YAML last-wins, harmless but confusing)
Wall-clock estimate ~13h on the existing setup. If results clearly
improve, consider 8× traversals (560) as a follow-up.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-07 16:40:32 +09:00
..
2026-05-07 16:32:44 +09:00
2026-05-07 16:32:44 +09:00
2026-05-07 16:32:44 +09:00
2026-05-07 16:40:32 +09:00
2026-05-07 16:32:44 +09:00
2026-05-07 16:32:44 +09:00
2026-05-07 16:32:44 +09:00
2026-05-07 16:32:44 +09:00
2026-05-07 16:32:44 +09:00
2026-05-07 16:32:44 +09:00
2026-05-07 16:32:44 +09:00
2026-05-07 16:32:44 +09:00
2026-05-07 16:32:44 +09:00
2026-05-07 16:32:44 +09:00