Document Deep CFR baseline analysis
This commit is contained in:
@@ -0,0 +1,207 @@
|
|||||||
|
# Deep CFR 512x3 2000-Iteration Baseline Analysis
|
||||||
|
|
||||||
|
**Last verified:** 2026-05-08, commit `aa898ed`
|
||||||
|
**Source run:** `runs/2026-05-08_051124_baseline-512x3-2000-dense-eval`
|
||||||
|
|
||||||
|
## Question
|
||||||
|
|
||||||
|
Does the 512x3 2000-iteration baseline show useful learning signals and
|
||||||
|
Lost Cities selectivity, and what should be improved before treating it as the
|
||||||
|
canonical comparison anchor?
|
||||||
|
|
||||||
|
## Short Answer
|
||||||
|
|
||||||
|
The run shows real optimization progress in the advantage networks, but the
|
||||||
|
policy that is evaluated does not improve. The main failure mode is not simply
|
||||||
|
"opens too many colors." By late training, the model often opens slightly fewer
|
||||||
|
colors against strong heuristic opponents, but the expeditions it opens are
|
||||||
|
lower quality and become negative more often. This looks more like passive or
|
||||||
|
degenerate average-policy learning than successful selectivity.
|
||||||
|
|
||||||
|
Use this run as evidence for the next diagnostics, not as a final baseline for
|
||||||
|
model-size decisions. Before running another long model-size comparison, separate
|
||||||
|
current regret-matching policy quality from average-strategy-network quality and
|
||||||
|
audit the interleaved outcome-sampling target behavior.
|
||||||
|
|
||||||
|
## Run Context
|
||||||
|
|
||||||
|
The source run used the old pre-deterministic baseline settings:
|
||||||
|
|
||||||
|
- `network.hidden_size=512`, `network.num_layers=3`
|
||||||
|
- `run.seed=79`
|
||||||
|
- `run.max_iterations=2000`
|
||||||
|
- `evaluation.eval_every=10`
|
||||||
|
- `checkpoint.save_every=100`
|
||||||
|
- `traversal.num_workers=8`
|
||||||
|
- `traversal.scheduler=interleaved`
|
||||||
|
- `traversal.sampling_mode=outcome`
|
||||||
|
- `traversal.outcome_unsampled_regret=zero`
|
||||||
|
- `run.deterministic=false`
|
||||||
|
|
||||||
|
It is therefore not the canonical reproducibility baseline after the deterministic
|
||||||
|
traversal changes. Its value is diagnostic: it is long enough to show how the
|
||||||
|
current 512x3 setup behaves after memory warm-up and many strategy updates.
|
||||||
|
|
||||||
|
## Evidence
|
||||||
|
|
||||||
|
Advantage fitting improves, but strategy fitting does not produce a better
|
||||||
|
evaluated policy.
|
||||||
|
|
||||||
|
| Metric | Early | Late | Interpretation |
|
||||||
|
| --- | ---: | ---: | --- |
|
||||||
|
| `loss/advantage` mean | ~2403 at iters 1-100 | ~807 at iters 1501-2000 | advantage networks fit their targets better |
|
||||||
|
| `loss/strategy` mean | ~1.53 at iters 1-100 | ~1.61 at iters 1501-2000 | average policy training does not clearly improve |
|
||||||
|
| `memory/advantage` | full by iter 101-500 | full | both player buffers saturated at 2M each |
|
||||||
|
| `memory/strategy` | full by iter 101-500 | full | strategy buffer saturated at 2M |
|
||||||
|
|
||||||
|
Evaluation performance peaks early and then degrades.
|
||||||
|
|
||||||
|
| Opponent | First 20 evals `win_rate0` | Last 20 evals `win_rate0` | First 20 evals score diff | Last 20 evals score diff |
|
||||||
|
| --- | ---: | ---: | ---: | ---: |
|
||||||
|
| random | 0.808 | 0.665 | +35.1 | +12.4 |
|
||||||
|
| passive_discard | 0.026 | 0.012 | -30.8 | -43.0 |
|
||||||
|
| safe_heuristic | 0.070 | 0.011 | -68.7 | -94.2 |
|
||||||
|
| safe_heuristic_loose | 0.078 | 0.014 | -69.3 | -95.0 |
|
||||||
|
| safe_heuristic_strict | 0.072 | 0.011 | -57.6 | -85.0 |
|
||||||
|
| noisy_safe | 0.120 | 0.033 | -49.4 | -78.8 |
|
||||||
|
|
||||||
|
The best strict-heuristic point appears around iteration 70:
|
||||||
|
|
||||||
|
- `eval/safe_heuristic_strict/win_rate0=0.15`
|
||||||
|
- `eval/safe_heuristic_strict/avg_score_diff0=-41.09`
|
||||||
|
|
||||||
|
The final point at iteration 2000 is worse:
|
||||||
|
|
||||||
|
- `eval/safe_heuristic_strict/win_rate0=0.02`
|
||||||
|
- `eval/safe_heuristic_strict/avg_score_diff0=-81.64`
|
||||||
|
|
||||||
|
Selectivity does not meaningfully emerge. Against `safe_heuristic_strict`:
|
||||||
|
|
||||||
|
| Metric | First 20 evals | Last 20 evals | Interpretation |
|
||||||
|
| --- | ---: | ---: | --- |
|
||||||
|
| `avg_opened_colors` | 4.31 | 3.92 | opens slightly fewer colors |
|
||||||
|
| `score_per_opened_color` | -8.09 | -15.41 | opened expeditions become worse |
|
||||||
|
| `positive_expedition_rate` | 0.142 | 0.046 | fewer opened expeditions end positive |
|
||||||
|
| `negative_expedition_rate` | 0.834 | 0.947 | almost every opened expedition is negative |
|
||||||
|
| `bad_open_rate` | 0.905 | 0.908 | first-open quality does not improve |
|
||||||
|
|
||||||
|
The open-quality metrics are generated by `src/coolrl_lost_cities/games/classic/deep_cfr/evaluate.py:124`
|
||||||
|
and classify first-open actions around `src/coolrl_lost_cities/games/classic/deep_cfr/evaluate.py:546`.
|
||||||
|
A "good" open has non-negative visible recoverable score; a "bad" open has no
|
||||||
|
bonus path and negative recoverable score.
|
||||||
|
|
||||||
|
Traversal gets deeper and slower late in training:
|
||||||
|
|
||||||
|
- `traversal/nodes` mean rises from ~119k at iters 1-100 to ~199k at iters
|
||||||
|
1501-2000.
|
||||||
|
- `traversal/avg_endpoint_depth` rises from ~212 to ~354.
|
||||||
|
- `time/iteration_seconds` rises from ~12.0s to ~14.4s.
|
||||||
|
- late node-limit cutoffs appear, but remain small on average: about 1.0% of
|
||||||
|
endpoints in iters 1501-2000.
|
||||||
|
|
||||||
|
This suggests the policy is changing game length and trajectory shape, but not
|
||||||
|
toward higher-quality scoring.
|
||||||
|
|
||||||
|
## Interpretation
|
||||||
|
|
||||||
|
The strongest signal is a current-vs-average policy split. Deep CFR trains
|
||||||
|
advantage networks to define a current regret-matching policy, but the evaluated
|
||||||
|
agent is the strategy network, which approximates the average policy. In this
|
||||||
|
run, advantage loss improves while strategy loss and eval quality do not. That
|
||||||
|
points to either:
|
||||||
|
|
||||||
|
1. the current policy is improving but the average strategy network or strategy
|
||||||
|
memory is failing to represent it, or
|
||||||
|
2. the advantage networks are fitting targets that do not translate into better
|
||||||
|
game play under the current traversal distribution.
|
||||||
|
|
||||||
|
The current metrics cannot distinguish those two cases. The next diagnostic
|
||||||
|
should evaluate the same checkpoint under both policy views:
|
||||||
|
|
||||||
|
- current advantage networks + regret matching,
|
||||||
|
- average strategy network.
|
||||||
|
|
||||||
|
There is also a target-consistency issue worth auditing before a new long run.
|
||||||
|
The recursive Cython traversal honors `outcome_unsampled_regret=zero` by setting
|
||||||
|
unsampled legal targets to zero at
|
||||||
|
`src/coolrl_lost_cities/games/classic/deep_cfr/traversal.pyx:930`. The
|
||||||
|
interleaved traversal currently sets every legal target to `-node_value` before
|
||||||
|
overwriting the sampled action at
|
||||||
|
`src/coolrl_lost_cities/games/classic/deep_cfr/interleaved_traversal.py:386`.
|
||||||
|
That makes the interleaved path behave like the `negative_node_value` variant
|
||||||
|
even when the resolved config says `zero`.
|
||||||
|
|
||||||
|
Theoretically, this is a state-wise constant shift and regret matching is
|
||||||
|
invariant to action-independent constants. With function approximation and MSE,
|
||||||
|
however, it changes training dynamics and should be treated as an experimental
|
||||||
|
variable unless explicitly made consistent.
|
||||||
|
|
||||||
|
## Deep CFR Improvement Candidates
|
||||||
|
|
||||||
|
### 1. Evaluate current policy vs average policy first
|
||||||
|
|
||||||
|
This is the highest-value diagnostic. If current regret-matching policy is
|
||||||
|
materially stronger than the strategy network, focus on strategy memory,
|
||||||
|
strategy weighting, and average-policy fitting. If both are weak, focus on
|
||||||
|
advantage targets, exploration, and representation.
|
||||||
|
|
||||||
|
### 2. Make interleaved outcome targets match the config
|
||||||
|
|
||||||
|
Either implement `outcome_unsampled_regret` in the interleaved scheduler or
|
||||||
|
change the default/config documentation to state that interleaved outcome
|
||||||
|
sampling uses the baseline-subtracted target. The former is cleaner because it
|
||||||
|
lets `zero` vs `negative_node_value` become a controlled ablation.
|
||||||
|
|
||||||
|
### 3. Run a lower-variance traversal ablation
|
||||||
|
|
||||||
|
Outcome sampling is correct but high variance for Lost Cities entry decisions.
|
||||||
|
Before changing model size, test whether increasing `traversals_per_player`
|
||||||
|
improves strict-heuristic score and expedition quality. A useful paired
|
||||||
|
ablation would keep model/config fixed and change only traversal count, for
|
||||||
|
example 280 vs 560 per player.
|
||||||
|
|
||||||
|
### 4. Add first-open regret diagnostics
|
||||||
|
|
||||||
|
The current bad-open metric is outcome-facing and heuristic-visible. It tells us
|
||||||
|
that opened expeditions are bad by recoverable score, but not whether the
|
||||||
|
counterfactual regret target said the open action was bad at the decision point.
|
||||||
|
Add diagnostics over advantage-memory targets by action class:
|
||||||
|
|
||||||
|
- first-open action targets,
|
||||||
|
- non-open alternatives in the same information state,
|
||||||
|
- color/action class buckets,
|
||||||
|
- target quantiles, not only means.
|
||||||
|
|
||||||
|
If bad first-open actions have negative advantage targets but the policy still
|
||||||
|
opens, suspect representation or strategy fitting. If their targets are not
|
||||||
|
negative, suspect exploration distribution or the metric/objective mismatch.
|
||||||
|
|
||||||
|
### 5. Treat model-size scaling as a follow-up, not the first fix
|
||||||
|
|
||||||
|
The earlier 1024x6 run showed a better 200-iteration signal than 512x3, but this
|
||||||
|
2000-iteration baseline shows a more basic policy-quality issue. Scaling capacity
|
||||||
|
may help representation, but it will not tell us whether the average-policy
|
||||||
|
training path is currently suppressing or distorting a better current policy.
|
||||||
|
|
||||||
|
## Recommended Next Plan
|
||||||
|
|
||||||
|
1. Implement or script current-vs-average evaluation for a saved checkpoint.
|
||||||
|
2. Run it on `iteration_00070.pt`, `iteration_00200.pt`, `iteration_01000.pt`,
|
||||||
|
and `iteration_02000.pt` from the source run.
|
||||||
|
3. Audit and fix interleaved `outcome_unsampled_regret` handling.
|
||||||
|
4. Establish a new deterministic 512x3 baseline only after the target behavior
|
||||||
|
is intentional and documented.
|
||||||
|
5. Then compare 512x3 vs 1024x6 under matched deterministic settings and W&B
|
||||||
|
grouping.
|
||||||
|
|
||||||
|
## Practical Baseline Rule
|
||||||
|
|
||||||
|
Do not promote `baseline-512x3-2000-dense-eval` as the canonical baseline. It is
|
||||||
|
better classified as a diagnostic pre-determinism long run. The next canonical
|
||||||
|
baseline should be rerun with:
|
||||||
|
|
||||||
|
- deterministic traversal enabled,
|
||||||
|
- `evaluation.eval_every=5`,
|
||||||
|
- `checkpoint.save_every=50`,
|
||||||
|
- explicit W&B group/job type,
|
||||||
|
- and target semantics that match the resolved config.
|
||||||
Reference in New Issue
Block a user