From bec59dfc3cc8acd437579d1f88dd2f7adc486202 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?=EC=A0=95=EC=8B=9C=EC=9B=90?= Date: Sun, 10 May 2026 22:54:03 +0900 Subject: [PATCH] =?UTF-8?q?Record=20=C2=A712=20R3=20SO-ISMCTS=20mini=20Los?= =?UTF-8?q?t=20Cities=20PoC=20result?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Mini Lost Cities (3색 5랭크) 50 iter / 20 eval games: - score_diff vs random +33, vs heuristic_cautious -5.7 (45% 승률) - play_action_rate 13~31% (Deep CFR의 0~2% 대비 명확) - trap escaped on mini, "long-horizon credit assignment" 가설 지지 Codex commit e69f316으로 ISMCTS 구현 완료. Full game scale-up이 다음 후보 (100 iter ETA 3시간 추정). --- docs/plans/deep-cfr-selectivity.md | 84 ++++++++++++++++++++++++++++++ 1 file changed, 84 insertions(+) diff --git a/docs/plans/deep-cfr-selectivity.md b/docs/plans/deep-cfr-selectivity.md index 36e954d..38ebd8f 100644 --- a/docs/plans/deep-cfr-selectivity.md +++ b/docs/plans/deep-cfr-selectivity.md @@ -980,3 +980,87 @@ R2의 advantage loss 200은 R0의 2200 대비 1/10. **모델이 trap 정책을 설계 문서 작성 + 단계별 구현은 별도 plan으로 (추가 예정). + +## §12. R3 — SO-ISMCTS mini Lost Cities PoC (2026-05-10) + +### Hypothesis + +R0/R1/R2 누적 진단 결과 trap 본질이 알고리즘 자체의 long-horizon credit +assignment 실패임이 확정. 다음 정공법으로 SO-ISMCTS (Single-Observer +Information Set MCTS) + AlphaZero-style policy/value head 도입. value head로 +terminal까지 walk 안 하고 search 부트스트랩으로 가치 추정. + +이번 실험은 **mini Lost Cities (3색 5랭크, hand 4, bonus 4)** 에서 알고리즘이 +작동하는지 PoC 검증. mini로 시작한 이유: full game은 horizon 길어 학습 비용 +큼, mini에서 알고리즘 작동 확인 후 scale-up. + +### Design + +- 변종: SO-ISMCTS (정보집합 노드 트리 + 시뮬마다 새 determinization) +- Determinization: uniform from unseen cards (Bayesian belief 미도입) +- Selection: PUCT (c_puct=1.5) +- Network: AlphaZero식 backbone (MLP 128×2 mini용) + policy head + tanh×100 value head +- Loss: MSE(value) + KL(policy) +- Tree: 매 결정마다 fresh (subtree 재사용 없음) +- Chance: closed-loop within sim, re-determinize across sims +- Pure self-play (양 플레이어 같은 네트워크) +- 50 simulations / decision (PoC, 향후 200+) + +### Implementation + +신규 모듈: `src/coolrl_lost_cities/games/classic/ismcts/` +- info_set.py, determinization.py, network.py, mcts.py, replay_buffer.py, + self_play.py, trainer.py, config.py, cli.py +- CLI: `lost-cities-ismcts train --config configs/ismcts/{mini,default}.yaml` +- 단위 테스트 7개 (정보집합 카노니컬, determinization 정합성, 네트워크 forward, + MCTS smoke, self-play 게임 종료, replay 버퍼) +- 기존 Deep CFR 코드 무수정 (encoding, evaluate만 재사용) + +Codex commit: `e69f316`. + +### Result (mini, 50 iter, 20 eval games per opponent) + +| 상대 | score_diff | win_rate | play_action_rate | +| --- | ---: | ---: | ---: | +| random | **+33.1** | 0.90 | 0.13 | +| **heuristic_cautious** | **-5.7** | **0.45** | **0.22** | +| discard_only | -20.8 | 0.00 | 0.31 | + +비교 (full game Deep CFR R0/R1/R2): + +| metric | Deep CFR (full game) | SO-ISMCTS (mini game) | +| --- | ---: | ---: | +| score_diff vs cautious | -48 ~ -70 | **-5.7** | +| heuristic_cautious 승률 | 0.06 | **0.45** | +| play_action_rate vs cautious | 0~2% | **22%** | + +### Interpretation: trap escaped on mini + +**Deep CFR이 6주 동안 못 풀던 핵심 결함 (play_action_rate ≈ 0%)이 단번에 +해소.** 모델이 처음으로 expedition에 카드를 의미 있게 두고 있음 (13~31%). +heuristic_cautious 상대 45% 승률은 거의 호각 (mini라 quality threshold 다르긴 +하지만 명확히 다른 학습 regime). + +진단 가치: **"trap = long-horizon credit assignment, search 부트스트랩으로 풀 +수 있음"** 가설 강하게 지지됨. mini는 horizon 짧아 (게임당 ~15-25 결정) +search bootstrap의 효과가 직접 확인 가능했음. + +### Caveats + +- Mini game은 credit chain이 짧음. full game (5색 9랭크, ~200 결정/게임)에선 + 같은 효과 보인다는 보장 없음. +- Eval 20게임 small sample. SE 큼. +- 시뮬 수 50은 mini에는 충분했지만 full은 부족할 수 있음. +- Strategy fusion / sparse exploration 같은 SO-ISMCTS 자체의 알려진 한계는 + mini에서 노출 안 됨. full에선 노출될 가능성. + +### Next: scale-up to full game + +R3a 후보: full Lost Cities (5색 9랭크) 100~200 iter 학습. ETA 3~6시간 추정 +(probe 측정 기반). configs/ismcts/default.yaml 기본 설정으로 시작. + +성공 기준 (mini와 평행): +- vs random: score_diff > +30, win_rate > 0.7 +- vs heuristic_cautious: play_action_rate > 10% 이상 (Deep CFR의 ~0%와 분명한 차이) +- 가능하면 score_diff > -30 (Deep CFR의 -50~-70보다 명확히 좋음) +