Files
coorl-lost-cities/docs/reports/final-cycles-and-human-play-2026-07-05.md
T
2026-07-14 20:09:03 +09:00

84 lines
3.7 KiB
Markdown

# Final Cycles and Human Play - 2026-07-05
## Part A - Closing Reinforcement Cycles
Raw artifacts: `/mnt/2tbhdd/coolrl-lost-cities-artifacts/final-cycles/2026-07-05`
Stop reason: `h2h_stagnation`
Final candidate: `/mnt/2tbhdd/coolrl-lost-cities-artifacts/final-cycles/2026-07-05/final_candidate`
Final candidate config: `/mnt/2tbhdd/coolrl-lost-cities-artifacts/final-cycles/2026-07-05/league/2026-07-05_191933_jax-ppo-final-cycle-c01/main_ppo_config.json`
The initial 4 hour GPU budget was treated as an estimate after user confirmation,
not a hard cutoff. The run completed cycle 1 fully, then stopped by the written
H2H stagnation rule. Cycle 2 was not started.
Bottom line: the closing reinforcement cycle did not produce a materially better
policy. The cycle-1 snapshot passed the expert guard, but the strengthened
exploiter battery got substantially worse than the incoming baseline target:
worst exploiter win rate was 0.7792, far above the 0.5200 success threshold. The
adjacent H2H against `repair_c01_update_500` was statistically indistinguishable
from zero, so this recipe is stopped here. Per instruction, this report does not
recommend more training with the same observation/network/PPO recipe.
### Expert Guard
| Cycle | Passed | Win rate | Mean diff | CI low | Opened colors | Max-step |
| ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| 1 | True | 0.7788 | +36.0847 | +34.6070 | 4.1523 | 0.0000 |
### Adjacent H2H
| Cycle | Current | Previous | Win rate | Mean diff | Score CI |
| ---: | --- | --- | ---: | ---: | ---: |
| 1 | `final_cycle_01` | `repair_c01_update_500` | 0.4905 | +0.3543 | [-1.0806, +1.7891] |
### Strengthened Exploiter Battery
| Cycle | Exploiter | Win rate | Mean diff | CI low | Opened colors | Max-step |
| ---: | --- | ---: | ---: | ---: | ---: | ---: |
| 1 | `long_random` | 0.7648 | +55.7990 | +53.6348 | 4.9242 | 0.0005 |
| 1 | `warmstart_gate3` | 0.7552 | +53.3092 | +51.0962 | 4.9008 | 0.0008 |
| 1 | `replay_exploiter` | 0.7792 | +59.3510 | +57.1215 | 4.9890 | 0.0003 |
### Battery Judgment
| Cycle | Worst exploiter | Worst win rate | Threshold | Passed |
| ---: | --- | ---: | ---: | ---: |
| 1 | `replay_exploiter` | 0.7792 | 0.5200 | False |
## Part B - Human Play Interface
Start a single game:
```bash
uv run --with 'jax[cuda12]' lost-cities-jax-ppo play \
--checkpoint /mnt/2tbhdd/coolrl-lost-cities-artifacts/final-cycles/2026-07-05/final_candidate \
--seat 0
```
Start a duplicate set with one shared shuffle and swapped seats:
```bash
uv run --with 'jax[cuda12]' lost-cities-jax-ppo play \
--checkpoint /mnt/2tbhdd/coolrl-lost-cities-artifacts/final-cycles/2026-07-05/final_candidate \
--seat 0 --duplicate
```
Summarize logged human games:
```bash
uv run lost-cities-jax-ppo human-play summarize \
--log-dir /mnt/2tbhdd/coolrl-lost-cities-artifacts/human-play/
```
Move syntax: `play R7 draw deck`, `discard G3 draw Y`, or `play RHS draw deck`.
The renderer shows only the human hand, both boards, all public discard piles, deck count, and current board score differential. Opponent hand and deck order are not rendered.
Every game is appended to `/mnt/2tbhdd/coolrl-lost-cities-artifacts/human-play/games.jsonl` with deck seed/index, full action list, AI top-3 policy actions/probabilities, value outputs, scoring breakdown, and optional human comment.
Validation performed before final report:
- `uv run ruff check .` passed.
- `uv run pytest -q` passed: 257 passed, 1 skipped.
- `uv run lost-cities-jax-ppo play --help` passed.
- `uv run lost-cities-jax-ppo human-play summarize --log-dir /tmp/nonexistent-human-play-log` passed.
- `scripts/librarian.sh` found no link or code-citation errors; it still exits non-zero on the known pre-existing `docs/plans/archive/deep-cfr-selectivity.md` 500-line soft cap.