From 2429d1e21019b51a10e42a7121ddfc019d05b142 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?=EC=A0=95=EC=8B=9C=EC=9B=90?= Date: Thu, 7 May 2026 21:37:44 +0900 Subject: [PATCH] Trim ideas.md to research-thread index + scratch space MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 진단 가설/개입/측정 섹션은 docs/research/lost_cities_selectivity.md 에 verbatim으로 보존되어 있어 ideas.md 쪽 사본은 stale 위험 + 중복 유지 비용만 남는다. ideas.md는 brainstorm 인덱스로 축소하고, research/ 의 세 thread를 단일 진입점으로 정리. 기존의 "All-negative fallback 가설 — 검증됨, 부분 풀림" 같은 stale 표현도 함께 사라진다 (research doc에선 여전히 open hypothesis로 다뤄지고 있음). Co-Authored-By: Claude Opus 4.7 (1M context) --- ideas.md | 60 ++++++-------------------------------------------------- 1 file changed, 6 insertions(+), 54 deletions(-) diff --git a/ideas.md b/ideas.md index e742b75..ccbb5a2 100644 --- a/ideas.md +++ b/ideas.md @@ -1,61 +1,13 @@ -# Lost Cities Deep CFR — Ideas 압축 정리 +# Lost Cities Deep CFR — Ideas + +짧은 brainstorm 인덱스. 자세한 분석/카탈로그는 `docs/research/`로 옮김. ## Active Research Threads -- [Lost Cities selectivity 가설 / 4-model 분석](docs/research/lost_cities_selectivity.md) +- [Lost Cities selectivity 가설 / 4-model 분석](docs/research/lost_cities_selectivity.md) — 진단 가설, 개입 아이디어, 측정 항목 카탈로그 포함. - [Julia 포팅 검토 — 단일/멀티스레드/ML 평가](docs/research/julia_port_evaluation.md) - [최적화 순서 / lever 의존성](docs/research/optimization_sequencing.md) -## 진단 가설들 +## Scratch -- Variance 가설 — Draw variance가 entry regret signal 압도. Reward smoothing 1순위. -- Architecture 가설 — Flat MLP가 action-conditional advantage 분리 못 함. -- Open and recover 가설 — 정책이 entry filter 대신 수습 학습 수렴. -- Action-local credit assignment 가설 — Skip color action 부재로 regret 분산. Bad_open metric 자체 mismatch 가능. -- Average / League inertia 가설 — Current policy는 selectivity 학습해도 average/league가 끌고 있음. -- All-negative fallback 가설 — Uniform fallback이 학습 dynamic hole. 검증됨, 부분 풀림. -- Self-play attractor 가설 — 5-color stable equilibrium. -- Lost Cities NE = 5-color 가설 — 가능성 낮지만 0 아님. - -## 개입 아이디어들 - -### Diagnostic (cheap, 정보량 큼) - -- Current vs Average vs League policy 분리 측정 -- Empirical r̃ by action class -- BR-to-current 진단 -- Tabular Lost Cities oracle -- 색별 opening rate, unopened color score - -### Architectural - -- Open-gate (latent skip color) -- Action-factorized scorer + per-action features -- Color permutation equivariance/augmentation -- Slot-shared encoder -- Dueling head -- Two-headed MLP - -### Training dynamics - -- ✅ Argmax_tiebreak fallback (적용됨, mechanism 작동) -- LCFR / DCFR weighting (reservoir inertia 직접 공격) -- Reward smoothing -- Open-action regret sign auxiliary -- First-open replay reweighting -- Type-balanced RM epsilon -- League weight 조정 (older 비중 감소) -- Memory discounting - -### Framework level - -- PSRO-style population - -## 새 measurement 항목 - -- Δ_open 기반 calibration metric -- Policy 분리 (current/average/league) opened_colors -- Fallback breakdown (rate, action 분포, opened_colors bucket, tie rate) -- Empirical r̃ by action class -- 색별 opening rate, unopened rate -- Avoided open penalty proxy +(짧은 새 아이디어가 떠오르면 여기에. 한 줄 이상으로 커지면 research/로 옮긴다.)