Files
coorl-lost-cities/docs/reports/verification-pass-2026-07-04.md
T

7.4 KiB

Verification Pass - 2026-07-04

Status: FAIL - stop before league self-play. Implementation start commit: bb52ef9. Primary checkpoint under audit: /mnt/2tbhdd/coolrl-lost-cities-artifacts/jax-ppo-static-opponents/2026-07-04_230150_jax-ppo-cautious/latest. Verification artifacts: /mnt/2tbhdd/coolrl-lost-cities-artifacts/verification/2026-07-04_phase1/.

Decision

Do not start snapshot-pool league self-play from the current ladder artifacts. The duplicate/Wilson protocol is reproducible, and the obvious seat-perspective bug was not found, but the cautious anchor is not a trustworthy permanent Elo anchor: it stalls heavily in mirror play and gives the trained policy a very clean win while the trained policy opens almost all colors.

Transcript Audit

Transcript file:

docs/reports/verification-pass-2026-07-04-transcripts.txt

Rendered protocol: 10 fixed shuffle-bank deck orders, duplicate seat-swapped, for 20 total games of gate-3 checkpoint vs heuristic_cautious.

Summary from the transcript dump:

Metric Value
Games 20
Agent mean score diff +144.500
Agent score diff min/max -93 / +230
Agent opened colors mean 5.000
Mean game length 159.200
Length min / p50 / p95 / max 63 / 109 / 400 / 400
Max-steps rate 0.150
Handshake play events 86
Cautious openings 78
Cautious low openings below rank 7 3
Cautious discards immediately playable by opponent 162

Interpretation: the sample confirms the concern behind this pass. The agent is not winning by a restrained 2-3 color expert pattern; it opens all 5 colors on average and still wins by a wide margin. The cautious policy is also leaking many immediately useful discard tops.

Protocol Recheck

The original ladder evaluation was audited in code. The eval path uses:

  • explicit shuffle banks generated from evaluation.shuffle_bank_seed;
  • duplicate seat-swapped loops over learner seat 0 and learner seat 1;
  • Wilson confidence intervals from aggregate wins.

The protocol itself is not the main failure. The failure is that the original gate summary omitted max_steps_rate and score-diff distribution canaries.

Recomputed gate-3 duplicate evaluation with distribution canaries:

/mnt/2tbhdd/coolrl-lost-cities-artifacts/verification/2026-07-04_phase1/gate3_reverification_distribution.json
Metric Value
Games 20,000
Win rate 0.95960
Wilson 95% [0.95678, 0.96224]
Mean score diff +142.9178
Score diff p05 / p50 / p95 +21 / +152 / +227
Opened colors/game 4.98965
Positive expeditions/game 4.09470
Play action rate 0.27639
Mean game length 185.3179
Game length p50 / p95 / max 115 / 400 / 400
Max-steps rate 0.24070

The 24.07% forced-end rate is a high-severity canary failure for using this gate as a clean league baseline.

Perspective And Mirror Tests

New tests added:

  • hand-authored heuristic behavior checks for weak unopened hands, strong balanced opens, and P1 own-board perspective;
  • static-policy duplicate mirror checks for all three heuristic policies;
  • gate-3 checkpoint duplicate self-mirror check, skipped on CPU-only runs and executed under CUDA JAX locally.

Commands run:

uv run pytest -q tests/lost_cities_jax/test_ppo_stack.py
uv run --with 'jax[cuda12]' pytest -q \
  tests/lost_cities_jax/test_ppo_stack.py::test_gate3_checkpoint_duplicate_self_mirror_score_diff_is_zero

Results:

  • CPU focused tests: 11 passed, 1 skipped.
  • CUDA checkpoint self-mirror test: 1 passed.

Mirror evaluation artifacts:

/mnt/2tbhdd/coolrl-lost-cities-artifacts/verification/2026-07-04_phase1/mirror_discard_only.json
/mnt/2tbhdd/coolrl-lost-cities-artifacts/verification/2026-07-04_phase1/mirror_heuristic_balanced.json
/mnt/2tbhdd/coolrl-lost-cities-artifacts/verification/2026-07-04_phase1/mirror_heuristic_cautious.json
/mnt/2tbhdd/coolrl-lost-cities-artifacts/verification/2026-07-04_phase1/mirror_gate3_checkpoint.json
Mirror policy Mean diff Win/loss symmetry Max-steps rate Play action rate Mean length
discard_only 0.0 exact 0.0000 0.0000 44.0000
heuristic_balanced 0.0 exact 0.4683 0.0422 231.1210
heuristic_cautious 0.0 exact 0.8956 0.0053 368.8435
gate-3 checkpoint 0.0 exact 0.0000 0.7839 45.4292

Interpretation: the broad P0/P1 perspective bug is unlikely. Duplicate mirror score differences cancel exactly. The serious issue is heuristic quality: heuristic_cautious mirror play almost never opens expeditions and reaches forced termination in 89.56% of games.

Exploiter Baseline

Frozen opponent: gate-3 checkpoint. New exploiter: random init PPO, same 8192 x 400 rollout shape, 250 updates, shaping annealed by the standard schedule.

Training run:

/mnt/2tbhdd/coolrl-lost-cities-artifacts/verification/2026-07-04_235948_jax-ppo-gate3-exploiter/

Final train-row canaries:

Metric Value
return_mean +0.24118
play_action_rate 0.71362
opened_colors_mean 4.91016
positive_expeditions_mean 2.07922
game_length_mean 49.12524
max_steps_rate 0.00000

Duplicate exploiter evaluation:

/mnt/2tbhdd/coolrl-lost-cities-artifacts/verification/2026-07-04_phase1/exploiter_vs_gate3_duplicate.json
Metric Value
Games 20,000
Exploiter win rate 0.67115
Wilson 95% [0.66461, 0.67763]
Mean score diff +24.1371
Score diff p05 / p50 / p95 -50 / +21 / +110
Opened colors/game 4.89710
Positive expeditions/game 2.12865
Play action rate 0.74612
Mean game length 48.0017
Max-steps rate 0.00000

This is now the measured exploitability baseline. It is not by itself a stop condition, but it reinforces that the current gate-3 checkpoint is not a finished robust policy.

Findings

Severity Finding Evidence Consequence
High heuristic_cautious is a stalling anchor, not a strong cautious baseline. Mirror max-steps rate 89.56%, play action rate 0.0053. Do not use it as a permanent Elo anchor without fixing/replacing it.
High Gate-3 result is contaminated by forced-end games and over-opening. Gate-3 max-steps rate 24.07%, opened colors/game 4.98965. The ladder PASS remains reproducible, but it should not become a league baseline.
Medium Cautious discards many immediately playable cards to the opponent. Transcript sample: 162 such discards in 20 games. The agent may be exploiting discard leakage rather than learning robust play.
Medium A same-budget PPO exploiter beats the gate-3 checkpoint. Exploiter duplicate win rate 67.115%, mean diff +24.1371. League success should reduce this number, but Phase 2 should wait for anchor repair.
Low No broad seat perspective bug found. Static and checkpoint duplicate mirrors have exact zero mean diff. Seat symmetry is not the likely explanation for the ladder result.

Stop Condition

Phase 1 did not pass cleanly. Per the work order, Phase 2 is not started.

Recommended next work:

  1. Replace or repair heuristic_balanced and heuristic_cautious so mirror play has near-zero forced-end rate and realistic play/open rates.
  2. Add max_steps_rate and game-length quantiles to every gate report.
  3. Rerun the static ladder gates with the repaired anchors before initializing a snapshot-pool league.