5.1 KiB
Lost Cities JAX PPO Snapshot-Pool League Spec
Status: ready for implementation after ladder v2.
Last updated: 2026-07-05.
Warm start: /mnt/2tbhdd/coolrl-lost-cities-artifacts/ladder-v2/2026-07-05_013223_jax-ppo-ladder-v2-expert/latest.
Preconditions
The static-opponent ladder v2 has passed:
discard_only: win rate 1.0000, mean score diff +182.0670.heuristic_balanced: win rate 0.9409, mean score diff +119.9250.heuristic_expert: win rate 0.8382, mean score diff +43.3099, score-diff CI95 lower bound +42.7065.
Canary warnings remain:
- Gate-3 opened colors are 4.5509/game, not the desired 2-3 range.
- The v2 gate-3 checkpoint is exploitable: a 250-update PPO exploiter reached win rate 0.83135 and mean score diff +54.7721 against it.
Anchor Semantics
Anchors are not certificates of strength. They serve two narrower purposes:
- Fix the Elo scale so curves remain comparable over time.
- Provide style diversity so the league does not train only against recent policy snapshots.
Strength certification is tracked on separate axes:
- Duplicate mean score difference against
heuristic_expert. - Exploiter win-rate trend under the fixed exploiter protocol.
- Later human play, once the automated diagnostics are stable.
Initial Pool
Permanent anchors are never removed:
- Static anchors:
discard_only,heuristic_balanced,heuristic_cautious,heuristic_expert. - Learned ladder anchors:
- Gate 1:
/mnt/2tbhdd/coolrl-lost-cities-artifacts/ladder-v2/2026-07-05_010429_jax-ppo-ladder-v2-discard-only/latest - Gate 2:
/mnt/2tbhdd/coolrl-lost-cities-artifacts/ladder-v2/2026-07-05_011827_jax-ppo-ladder-v2-balanced/latest - Gate 3:
/mnt/2tbhdd/coolrl-lost-cities-artifacts/ladder-v2/2026-07-05_013223_jax-ppo-ladder-v2-expert/latest
- Gate 1:
The learning policy starts from the v2 gate-3 checkpoint. Shaping coefficient is fixed at 0 by default. If warm-start collapse is observed, the implementation may expose a config switch for random initialization plus shaping anneal, but that is not the default path.
Opponent Sampling
For each episode:
- Randomize the learner seat.
- With probability 50%, play mirror self-play against the current policy copy.
- With probability 50%, sample from the pool.
Pool sampling is PFSP-like:
- Track recent win rate versus each pool member.
- Weight non-anchor snapshots by
(1 - win_rate)^2. - Mix in a small uniform component to avoid starvation.
- Give every permanent anchor only a small floor probability, such as 2-3% per anchor, so stalling or weak anchors do not waste most throughput.
- Allocate the remaining probability mass to snapshots according to the PFSP weights.
Snapshot Lifecycle
- Add the current policy to the pool at a configurable interval.
- If the pool size cap is reached, remove the oldest non-anchor snapshot.
- Never remove permanent anchors.
- Store large checkpoints and raw evaluation JSON under
/mnt/2tbhdd/coolrl-lost-cities-artifacts/league/. - Store one-line tracked JSON summaries under
docs/reports/so results do not disappear with artifact cleanup.
Evaluation
At each snapshot interval, evaluate the current policy with fixed shuffle-bank duplicate play against:
- The six ladder anchors.
heuristic_expert.- A configurable set of recent snapshots.
Update logistic Elo from the pairwise results, using the heuristic anchors as fixed reference points. Elo is a scale diagnostic, not a strength certificate.
Run the exploiter protocol periodically against the current best checkpoint:
- Randomly initialized PPO.
- Same 250-update budget used in ladder v2 unless explicitly overridden.
- Duplicate evaluation versus the frozen target checkpoint.
- Track exploiter win rate and mean score difference over time.
Canary Metrics
Log these metrics at every evaluation:
- Opened colors per game.
- Play action rate.
- Mean and quantile game length.
- Max-steps rate.
- Positive expeditions per game.
- Duplicate score-diff distribution against
heuristic_expert.
Canaries are diagnostics only. Do not put them directly into the reward.
Regression Guards
heuristic_expertreplacesheuristic_cautiousas the main regression opponent.- If win rate versus
heuristic_expertfalls below 90% of the v2 gate-3 value, flag the snapshot. - If max-steps rate rises above the gate-3 baseline of 0.0 by a material amount, flag the snapshot.
- If exploiter win rate rises above the v2 baseline of 0.83135, flag the snapshot.
Stop Conditions
Stop the league run when either:
- Elo is statistically flat over a configured recent-snapshot window and exploiter win rate is at or below the configured threshold, for example 0.55.
- The wall-clock budget is exhausted.
Decisions
- The earlier
heuristic_cautiousremains a permanent anchor for scale and style diversity only. It is a stalling anchor and no longer certifies strength. discard_onlyis also kept only as a low-probability anchor and regression sanity check.- The league spec file was not present in the repository when this update was requested, so this tracked plan file is the authoritative spec location.