Files
coorl-lost-cities/docs/research/deep-cfr-baseline-2000-analysis.md
T

208 lines
9.2 KiB
Markdown

# Deep CFR 512x3 2000-Iteration Baseline Analysis
**Last verified:** 2026-05-08, commit `aa898ed`
**Source run:** `runs/2026-05-08_051124_baseline-512x3-2000-dense-eval`
## Question
Does the 512x3 2000-iteration baseline show useful learning signals and
Lost Cities selectivity, and what should be improved before treating it as the
canonical comparison anchor?
## Short Answer
The run shows real optimization progress in the advantage networks, but the
policy that is evaluated does not improve. The main failure mode is not simply
"opens too many colors." By late training, the model often opens slightly fewer
colors against strong heuristic opponents, but the expeditions it opens are
lower quality and become negative more often. This looks more like passive or
degenerate average-policy learning than successful selectivity.
Use this run as evidence for the next diagnostics, not as a final baseline for
model-size decisions. Before running another long model-size comparison, separate
current regret-matching policy quality from average-strategy-network quality and
audit the interleaved outcome-sampling target behavior.
## Run Context
The source run used the old pre-deterministic baseline settings:
- `network.hidden_size=512`, `network.num_layers=3`
- `run.seed=79`
- `run.max_iterations=2000`
- `evaluation.eval_every=10`
- `checkpoint.save_every=100`
- `traversal.num_workers=8`
- `traversal.scheduler=interleaved`
- `traversal.sampling_mode=outcome`
- `traversal.outcome_unsampled_regret=zero`
- `run.deterministic=false`
It is therefore not the canonical reproducibility baseline after the deterministic
traversal changes. Its value is diagnostic: it is long enough to show how the
current 512x3 setup behaves after memory warm-up and many strategy updates.
## Evidence
Advantage fitting improves, but strategy fitting does not produce a better
evaluated policy.
| Metric | Early | Late | Interpretation |
| --- | ---: | ---: | --- |
| `loss/advantage` mean | ~2403 at iters 1-100 | ~807 at iters 1501-2000 | advantage networks fit their targets better |
| `loss/strategy` mean | ~1.53 at iters 1-100 | ~1.61 at iters 1501-2000 | average policy training does not clearly improve |
| `memory/advantage` | full by iter 101-500 | full | both player buffers saturated at 2M each |
| `memory/strategy` | full by iter 101-500 | full | strategy buffer saturated at 2M |
Evaluation performance peaks early and then degrades.
| Opponent | First 20 evals `win_rate0` | Last 20 evals `win_rate0` | First 20 evals score diff | Last 20 evals score diff |
| --- | ---: | ---: | ---: | ---: |
| random | 0.808 | 0.665 | +35.1 | +12.4 |
| passive_discard | 0.026 | 0.012 | -30.8 | -43.0 |
| safe_heuristic | 0.070 | 0.011 | -68.7 | -94.2 |
| safe_heuristic_loose | 0.078 | 0.014 | -69.3 | -95.0 |
| safe_heuristic_strict | 0.072 | 0.011 | -57.6 | -85.0 |
| noisy_safe | 0.120 | 0.033 | -49.4 | -78.8 |
The best strict-heuristic point appears around iteration 70:
- `eval/safe_heuristic_strict/win_rate0=0.15`
- `eval/safe_heuristic_strict/avg_score_diff0=-41.09`
The final point at iteration 2000 is worse:
- `eval/safe_heuristic_strict/win_rate0=0.02`
- `eval/safe_heuristic_strict/avg_score_diff0=-81.64`
Selectivity does not meaningfully emerge. Against `safe_heuristic_strict`:
| Metric | First 20 evals | Last 20 evals | Interpretation |
| --- | ---: | ---: | --- |
| `avg_opened_colors` | 4.31 | 3.92 | opens slightly fewer colors |
| `score_per_opened_color` | -8.09 | -15.41 | opened expeditions become worse |
| `positive_expedition_rate` | 0.142 | 0.046 | fewer opened expeditions end positive |
| `negative_expedition_rate` | 0.834 | 0.947 | almost every opened expedition is negative |
| `bad_open_rate` | 0.905 | 0.908 | first-open quality does not improve |
The open-quality metrics are generated by `src/coolrl_lost_cities/games/classic/deep_cfr/evaluate.py:124`
and classify first-open actions around `src/coolrl_lost_cities/games/classic/deep_cfr/evaluate.py:546`.
A "good" open has non-negative visible recoverable score; a "bad" open has no
bonus path and negative recoverable score.
Traversal gets deeper and slower late in training:
- `traversal/nodes` mean rises from ~119k at iters 1-100 to ~199k at iters
1501-2000.
- `traversal/avg_endpoint_depth` rises from ~212 to ~354.
- `time/iteration_seconds` rises from ~12.0s to ~14.4s.
- late node-limit cutoffs appear, but remain small on average: about 1.0% of
endpoints in iters 1501-2000.
This suggests the policy is changing game length and trajectory shape, but not
toward higher-quality scoring.
## Interpretation
The strongest signal is a current-vs-average policy split. Deep CFR trains
advantage networks to define a current regret-matching policy, but the evaluated
agent is the strategy network, which approximates the average policy. In this
run, advantage loss improves while strategy loss and eval quality do not. That
points to either:
1. the current policy is improving but the average strategy network or strategy
memory is failing to represent it, or
2. the advantage networks are fitting targets that do not translate into better
game play under the current traversal distribution.
The current metrics cannot distinguish those two cases. The next diagnostic
should evaluate the same checkpoint under both policy views:
- current advantage networks + regret matching,
- average strategy network.
There is also a target-consistency issue worth auditing before a new long run.
The recursive Cython traversal honors `outcome_unsampled_regret=zero` by setting
unsampled legal targets to zero at
`src/coolrl_lost_cities/games/classic/deep_cfr/traversal.pyx:930`. The
interleaved traversal currently sets every legal target to `-node_value` before
overwriting the sampled action at
`src/coolrl_lost_cities/games/classic/deep_cfr/interleaved_traversal.py:386`.
That makes the interleaved path behave like the `negative_node_value` variant
even when the resolved config says `zero`.
Theoretically, this is a state-wise constant shift and regret matching is
invariant to action-independent constants. With function approximation and MSE,
however, it changes training dynamics and should be treated as an experimental
variable unless explicitly made consistent.
## Deep CFR Improvement Candidates
### 1. Evaluate current policy vs average policy first
This is the highest-value diagnostic. If current regret-matching policy is
materially stronger than the strategy network, focus on strategy memory,
strategy weighting, and average-policy fitting. If both are weak, focus on
advantage targets, exploration, and representation.
### 2. Make interleaved outcome targets match the config
Either implement `outcome_unsampled_regret` in the interleaved scheduler or
change the default/config documentation to state that interleaved outcome
sampling uses the baseline-subtracted target. The former is cleaner because it
lets `zero` vs `negative_node_value` become a controlled ablation.
### 3. Run a lower-variance traversal ablation
Outcome sampling is correct but high variance for Lost Cities entry decisions.
Before changing model size, test whether increasing `traversals_per_player`
improves strict-heuristic score and expedition quality. A useful paired
ablation would keep model/config fixed and change only traversal count, for
example 280 vs 560 per player.
### 4. Add first-open regret diagnostics
The current bad-open metric is outcome-facing and heuristic-visible. It tells us
that opened expeditions are bad by recoverable score, but not whether the
counterfactual regret target said the open action was bad at the decision point.
Add diagnostics over advantage-memory targets by action class:
- first-open action targets,
- non-open alternatives in the same information state,
- color/action class buckets,
- target quantiles, not only means.
If bad first-open actions have negative advantage targets but the policy still
opens, suspect representation or strategy fitting. If their targets are not
negative, suspect exploration distribution or the metric/objective mismatch.
### 5. Treat model-size scaling as a follow-up, not the first fix
The earlier 1024x6 run showed a better 200-iteration signal than 512x3, but this
2000-iteration baseline shows a more basic policy-quality issue. Scaling capacity
may help representation, but it will not tell us whether the average-policy
training path is currently suppressing or distorting a better current policy.
## Recommended Next Plan
1. Implement or script current-vs-average evaluation for a saved checkpoint.
2. Run it on `iteration_00070.pt`, `iteration_00200.pt`, `iteration_01000.pt`,
and `iteration_02000.pt` from the source run.
3. Audit and fix interleaved `outcome_unsampled_regret` handling.
4. Establish a new deterministic 512x3 baseline only after the target behavior
is intentional and documented.
5. Then compare 512x3 vs 1024x6 under matched deterministic settings and W&B
grouping.
## Practical Baseline Rule
Do not promote `baseline-512x3-2000-dense-eval` as the canonical baseline. It is
better classified as a diagnostic pre-determinism long run. The next canonical
baseline should be rerun with:
- deterministic traversal enabled,
- `evaluation.eval_every=5`,
- `checkpoint.save_every=50`,
- explicit W&B group/job type,
- and target semantics that match the resolved config.