Add Lost Cities selectivity research notes (4-model analysis)

Mechanism-level analysis of the slot_aware_playability iter 240
plateau (opened_colors 4.95+, bad_open_rate 88-92%, calibration
gap 6-9 → 2-4). Captures four expert consultations with diagnostic
hypotheses, intervention catalog (architectural / training-dynamics
/ game-specific), measurement plan, and a comparison table across
the four sources.

Key new directions surfaced:
- Current vs average vs league policy separation (Deep CFR average
  strategy is the convergence target, not advantage current).
- All-negative fallback as Deep CFR ablation lever.
- Empirical r̃ partitioning by action class.
- Tabular Lost Cities oracle as a clean test of "is 5-color the
  game-theoretic answer or an approximation artifact".
- Entry-gate target defined from traversal counterfactual values
  instead of heuristic labels.

ideas.md gets an "Active Research Threads" pointer.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-05-07 20:58:09 +09:00
co-authored by Claude Opus 4.7
parent 6c9babe769
commit 8f6780dd3d
2 changed files with 205 additions and 0 deletions
+4
View File
@@ -1,5 +1,9 @@
# Lost Cities Deep CFR — Ideas 압축 정리
## Active Research Threads
- [Lost Cities selectivity 가설 / 4-model 분석](docs/research/lost_cities_selectivity.md)
## 진단 가설들
- Variance 가설 — Draw variance가 entry regret signal 압도. Reward smoothing 1순위.