이전 레포 commit 279d726, a77464b의 self_play_league에 safe_heuristic anchor 0.15 주입 실험(1219 iter / 4h 풀 런) 결과를 doc에 기록. opened_colors 4.83 / 5-color 86%로 trap 못 깸. 0.15 weight으로는 self-mirror 평형 절단 불충분이라는 결론 명시.
이전 레포 commit 279d726, a77464b의 self_play_league에 safe_heuristic anchor 0.15 주입 실험(1219 iter / 4h 풀 런) 결과를 doc에 기록. opened_colors 4.83 / 5-color 86%로 trap 못 깸. 0.15 weight으로는 self-mirror 평형 절단 불충분이라는 결론 명시.