3.8 KiB
3.8 KiB
Gates 1-2 Report - 2026-07-05
Run dir: /mnt/2tbhdd/coolrl-lost-cities-artifacts/gates-1-2/2026-07-05_094238_gates-1-2.
Target: /mnt/2tbhdd/coolrl-lost-cities-artifacts/league/2026-07-05_052325_jax-ppo-league-v1/snapshots/cycle_01_update_000500.
Gate 1A - Policy Class Tournament
| Learner | Opponent | Win rate | Mean diff | CI low | Opened colors | Max-step |
|---|---|---|---|---|---|---|
expert_cap2 |
expert_cap3 |
0.4415 | -1.8933 | -2.3662 | 1.7792 | 0.0000 |
expert_cap2 |
expert_capN |
0.4445 | -1.8807 | -2.3545 | 1.7815 | 0.0000 |
expert_cap3 |
expert_capN |
0.4840 | +0.0285 | -0.4689 | 2.2153 | 0.0000 |
Gate 1A - Variants vs League v1
| Learner | Opponent | Win rate | Mean diff | CI low | Opened colors | Max-step |
|---|---|---|---|---|---|---|
expert_cap2 |
league_v1_update_500 |
0.2725 | -20.8863 | -21.9846 | 1.9735 | 0.0003 |
expert_cap3 |
league_v1_update_500 |
0.3400 | -14.3513 | -15.4459 | 2.7385 | 0.0000 |
expert_capN |
league_v1_update_500 |
0.3415 | -13.9670 | -15.0543 | 3.0040 | 0.0000 |
Gate 1B - Delta Open Audit
- States: 500
- Paired samples: 32000
- Mean delta: +0.0138
- CI95: [-0.3283, +0.3559]
- Judgment:
near_zero - Histogram:
docs/reports/gates-1-2-delta-open-hist.png
Gate 1C - Selectivity Judgment
통념 기각/미결: capN은 집중 변형보다 유의하게 나쁘지 않고, 리그 정책의 4번째+ 오픈 delta는 0 근처다. selectivity는 현재 주요 성능 병목으로 보이지 않는다.
Gate 2A - Strengthened Exploiter Battery
| Exploiter | Win rate | Mean diff | CI low | Opened colors | Max-step | Run |
|---|---|---|---|---|---|---|
long_random |
0.5870 | +10.9512 | +9.4790 | 4.9958 | 0.0000 | /mnt/2tbhdd/coolrl-lost-cities-artifacts/gates-1-2/2026-07-05_094238_gates-1-2/exploiters/2026-07-05_094311_gates-1-2-long-random-exploiter |
warmstart_gate3 |
0.5900 | +11.3808 | +9.8893 | 4.9945 | 0.0000 | /mnt/2tbhdd/coolrl-lost-cities-artifacts/gates-1-2/2026-07-05_094238_gates-1-2/exploiters/2026-07-05_105902_gates-1-2-warmstart-gate3-exploiter |
replay_exploiter |
0.5820 | +10.4253 | +8.9368 | 4.9977 | 0.0000 | /mnt/2tbhdd/coolrl-lost-cities-artifacts/gates-1-2/2026-07-05_094238_gates-1-2/exploiters/2026-07-05_115602_gates-1-2-replay-exploiter |
Gate 2B - Robustness Judgment
보수 필요: worst exploiter warmstart_gate3 win rate 0.5900 vs threshold 0.5500.
Gate 2C - Conditional Repair League
| Cycle | Worst protocol used | Guard pass | Battery worst | Passed | Checkpoint |
|---|---|---|---|---|---|
| 1 | warmstart_gate3 |
True | 0.5490 | True | /mnt/2tbhdd/coolrl-lost-cities-artifacts/gates-1-2/2026-07-05_094238_gates-1-2/gate2c/league/2026-07-05_125709_jax-ppo-gates-1-2-repair-c01/snapshots/cycle_01_update_000500 |
| Cycle | Exploiter | Win rate | Mean diff | CI low | Opened colors | Max-step |
|---|---|---|---|---|---|---|
| 1 | long_random |
0.5390 | +6.2793 | +4.7919 | 4.9943 | 0.0000 |
| 1 | warmstart_gate3 |
0.5490 | +7.0285 | +5.5662 | 4.9935 | 0.0000 |
| 1 | replay_exploiter |
0.5393 | +6.1022 | +4.6214 | 4.9943 | 0.0000 |
Human Play Recommendation
조건부 예: 강화 exploiter 관문은 통과했다. 다만 selectivity 관문이 미결/부분 지지이면 인간 대전은 실력 인증이 아니라 행동 양식 진단으로 시작해야 한다.
Decisions
- Existing
heuristic_expertremains unchanged. The cap variants add only a hardmax_open_colorsgate around new-color openings. expert_capNmeans no hard cap; EV thresholds and the existing soft concentration penalties are retained.- Warm-started exploiters use shaping coefficient 0 to measure target-specific exploitation without reintroducing early shaping rewards.