7.4 KiB
Verification Pass - 2026-07-04
Status: FAIL - stop before league self-play.
Implementation start commit: bb52ef9.
Primary checkpoint under audit: /mnt/2tbhdd/coolrl-lost-cities-artifacts/jax-ppo-static-opponents/2026-07-04_230150_jax-ppo-cautious/latest.
Verification artifacts: /mnt/2tbhdd/coolrl-lost-cities-artifacts/verification/2026-07-04_phase1/.
Decision
Do not start snapshot-pool league self-play from the current ladder artifacts. The duplicate/Wilson protocol is reproducible, and the obvious seat-perspective bug was not found, but the cautious anchor is not a trustworthy permanent Elo anchor: it stalls heavily in mirror play and gives the trained policy a very clean win while the trained policy opens almost all colors.
Transcript Audit
Transcript file:
docs/reports/verification-pass-2026-07-04-transcripts.txt
Rendered protocol: 10 fixed shuffle-bank deck orders, duplicate seat-swapped,
for 20 total games of gate-3 checkpoint vs heuristic_cautious.
Summary from the transcript dump:
| Metric | Value |
|---|---|
| Games | 20 |
| Agent mean score diff | +144.500 |
| Agent score diff min/max | -93 / +230 |
| Agent opened colors mean | 5.000 |
| Mean game length | 159.200 |
| Length min / p50 / p95 / max | 63 / 109 / 400 / 400 |
| Max-steps rate | 0.150 |
| Handshake play events | 86 |
| Cautious openings | 78 |
| Cautious low openings below rank 7 | 3 |
| Cautious discards immediately playable by opponent | 162 |
Interpretation: the sample confirms the concern behind this pass. The agent is not winning by a restrained 2-3 color expert pattern; it opens all 5 colors on average and still wins by a wide margin. The cautious policy is also leaking many immediately useful discard tops.
Protocol Recheck
The original ladder evaluation was audited in code. The eval path uses:
- explicit shuffle banks generated from
evaluation.shuffle_bank_seed; - duplicate seat-swapped loops over learner seat 0 and learner seat 1;
- Wilson confidence intervals from aggregate wins.
The protocol itself is not the main failure. The failure is that the original
gate summary omitted max_steps_rate and score-diff distribution canaries.
Recomputed gate-3 duplicate evaluation with distribution canaries:
/mnt/2tbhdd/coolrl-lost-cities-artifacts/verification/2026-07-04_phase1/gate3_reverification_distribution.json
| Metric | Value |
|---|---|
| Games | 20,000 |
| Win rate | 0.95960 |
| Wilson 95% | [0.95678, 0.96224] |
| Mean score diff | +142.9178 |
| Score diff p05 / p50 / p95 | +21 / +152 / +227 |
| Opened colors/game | 4.98965 |
| Positive expeditions/game | 4.09470 |
| Play action rate | 0.27639 |
| Mean game length | 185.3179 |
| Game length p50 / p95 / max | 115 / 400 / 400 |
| Max-steps rate | 0.24070 |
The 24.07% forced-end rate is a high-severity canary failure for using this gate as a clean league baseline.
Perspective And Mirror Tests
New tests added:
- hand-authored heuristic behavior checks for weak unopened hands, strong balanced opens, and P1 own-board perspective;
- static-policy duplicate mirror checks for all three heuristic policies;
- gate-3 checkpoint duplicate self-mirror check, skipped on CPU-only runs and executed under CUDA JAX locally.
Commands run:
uv run pytest -q tests/lost_cities_jax/test_ppo_stack.py
uv run --with 'jax[cuda12]' pytest -q \
tests/lost_cities_jax/test_ppo_stack.py::test_gate3_checkpoint_duplicate_self_mirror_score_diff_is_zero
Results:
- CPU focused tests: 11 passed, 1 skipped.
- CUDA checkpoint self-mirror test: 1 passed.
Mirror evaluation artifacts:
/mnt/2tbhdd/coolrl-lost-cities-artifacts/verification/2026-07-04_phase1/mirror_discard_only.json
/mnt/2tbhdd/coolrl-lost-cities-artifacts/verification/2026-07-04_phase1/mirror_heuristic_balanced.json
/mnt/2tbhdd/coolrl-lost-cities-artifacts/verification/2026-07-04_phase1/mirror_heuristic_cautious.json
/mnt/2tbhdd/coolrl-lost-cities-artifacts/verification/2026-07-04_phase1/mirror_gate3_checkpoint.json
| Mirror policy | Mean diff | Win/loss symmetry | Max-steps rate | Play action rate | Mean length |
|---|---|---|---|---|---|
discard_only |
0.0 | exact | 0.0000 | 0.0000 | 44.0000 |
heuristic_balanced |
0.0 | exact | 0.4683 | 0.0422 | 231.1210 |
heuristic_cautious |
0.0 | exact | 0.8956 | 0.0053 | 368.8435 |
| gate-3 checkpoint | 0.0 | exact | 0.0000 | 0.7839 | 45.4292 |
Interpretation: the broad P0/P1 perspective bug is unlikely. Duplicate mirror
score differences cancel exactly. The serious issue is heuristic quality:
heuristic_cautious mirror play almost never opens expeditions and reaches
forced termination in 89.56% of games.
Exploiter Baseline
Frozen opponent: gate-3 checkpoint. New exploiter: random init PPO, same 8192 x 400 rollout shape, 250 updates, shaping annealed by the standard schedule.
Training run:
/mnt/2tbhdd/coolrl-lost-cities-artifacts/verification/2026-07-04_235948_jax-ppo-gate3-exploiter/
Final train-row canaries:
| Metric | Value |
|---|---|
| return_mean | +0.24118 |
| play_action_rate | 0.71362 |
| opened_colors_mean | 4.91016 |
| positive_expeditions_mean | 2.07922 |
| game_length_mean | 49.12524 |
| max_steps_rate | 0.00000 |
Duplicate exploiter evaluation:
/mnt/2tbhdd/coolrl-lost-cities-artifacts/verification/2026-07-04_phase1/exploiter_vs_gate3_duplicate.json
| Metric | Value |
|---|---|
| Games | 20,000 |
| Exploiter win rate | 0.67115 |
| Wilson 95% | [0.66461, 0.67763] |
| Mean score diff | +24.1371 |
| Score diff p05 / p50 / p95 | -50 / +21 / +110 |
| Opened colors/game | 4.89710 |
| Positive expeditions/game | 2.12865 |
| Play action rate | 0.74612 |
| Mean game length | 48.0017 |
| Max-steps rate | 0.00000 |
This is now the measured exploitability baseline. It is not by itself a stop condition, but it reinforces that the current gate-3 checkpoint is not a finished robust policy.
Findings
| Severity | Finding | Evidence | Consequence |
|---|---|---|---|
| High | heuristic_cautious is a stalling anchor, not a strong cautious baseline. |
Mirror max-steps rate 89.56%, play action rate 0.0053. | Do not use it as a permanent Elo anchor without fixing/replacing it. |
| High | Gate-3 result is contaminated by forced-end games and over-opening. | Gate-3 max-steps rate 24.07%, opened colors/game 4.98965. | The ladder PASS remains reproducible, but it should not become a league baseline. |
| Medium | Cautious discards many immediately playable cards to the opponent. | Transcript sample: 162 such discards in 20 games. | The agent may be exploiting discard leakage rather than learning robust play. |
| Medium | A same-budget PPO exploiter beats the gate-3 checkpoint. | Exploiter duplicate win rate 67.115%, mean diff +24.1371. | League success should reduce this number, but Phase 2 should wait for anchor repair. |
| Low | No broad seat perspective bug found. | Static and checkpoint duplicate mirrors have exact zero mean diff. | Seat symmetry is not the likely explanation for the ladder result. |
Stop Condition
Phase 1 did not pass cleanly. Per the work order, Phase 2 is not started.
Recommended next work:
- Replace or repair
heuristic_balancedandheuristic_cautiousso mirror play has near-zero forced-end rate and realistic play/open rates. - Add
max_steps_rateand game-length quantiles to every gate report. - Rerun the static ladder gates with the repaired anchors before initializing a snapshot-pool league.