# Deep CFR 512x3 2000-Iteration Baseline Analysis **Last verified:** 2026-05-08, commit `aa898ed` **Source run:** `runs/2026-05-08_051124_baseline-512x3-2000-dense-eval` ## Question Does the 512x3 2000-iteration baseline show useful learning signals and Lost Cities selectivity, and what should be improved before treating it as the canonical comparison anchor? ## Short Answer The run shows real optimization progress in the advantage networks, but the policy that is evaluated does not improve. The main failure mode is not simply "opens too many colors." By late training, the model often opens slightly fewer colors against strong heuristic opponents, but the expeditions it opens are lower quality and become negative more often. This looks more like passive or degenerate average-policy learning than successful selectivity. Use this run as evidence for the next diagnostics, not as a final baseline for model-size decisions. Before running another long model-size comparison, separate current regret-matching policy quality from average-strategy-network quality and audit the interleaved outcome-sampling target behavior. ## Run Context The source run used the old pre-deterministic baseline settings: - `network.hidden_size=512`, `network.num_layers=3` - `run.seed=79` - `run.max_iterations=2000` - `evaluation.eval_every=10` - `checkpoint.save_every=100` - `traversal.num_workers=8` - `traversal.scheduler=interleaved` - `traversal.sampling_mode=outcome` - `traversal.outcome_unsampled_regret=zero` - `run.deterministic=false` It is therefore not the canonical reproducibility baseline after the deterministic traversal changes. Its value is diagnostic: it is long enough to show how the current 512x3 setup behaves after memory warm-up and many strategy updates. ## Evidence Advantage fitting improves, but strategy fitting does not produce a better evaluated policy. | Metric | Early | Late | Interpretation | | --- | ---: | ---: | --- | | `loss/advantage` mean | ~2403 at iters 1-100 | ~807 at iters 1501-2000 | advantage networks fit their targets better | | `loss/strategy` mean | ~1.53 at iters 1-100 | ~1.61 at iters 1501-2000 | average policy training does not clearly improve | | `memory/advantage` | full by iter 101-500 | full | both player buffers saturated at 2M each | | `memory/strategy` | full by iter 101-500 | full | strategy buffer saturated at 2M | Evaluation performance peaks early and then degrades. | Opponent | First 20 evals `win_rate0` | Last 20 evals `win_rate0` | First 20 evals score diff | Last 20 evals score diff | | --- | ---: | ---: | ---: | ---: | | random | 0.808 | 0.665 | +35.1 | +12.4 | | passive_discard | 0.026 | 0.012 | -30.8 | -43.0 | | safe_heuristic | 0.070 | 0.011 | -68.7 | -94.2 | | safe_heuristic_loose | 0.078 | 0.014 | -69.3 | -95.0 | | safe_heuristic_strict | 0.072 | 0.011 | -57.6 | -85.0 | | noisy_safe | 0.120 | 0.033 | -49.4 | -78.8 | The best strict-heuristic point appears around iteration 70: - `eval/safe_heuristic_strict/win_rate0=0.15` - `eval/safe_heuristic_strict/avg_score_diff0=-41.09` The final point at iteration 2000 is worse: - `eval/safe_heuristic_strict/win_rate0=0.02` - `eval/safe_heuristic_strict/avg_score_diff0=-81.64` Selectivity does not meaningfully emerge. Against `safe_heuristic_strict`: | Metric | First 20 evals | Last 20 evals | Interpretation | | --- | ---: | ---: | --- | | `avg_opened_colors` | 4.31 | 3.92 | opens slightly fewer colors | | `score_per_opened_color` | -8.09 | -15.41 | opened expeditions become worse | | `positive_expedition_rate` | 0.142 | 0.046 | fewer opened expeditions end positive | | `negative_expedition_rate` | 0.834 | 0.947 | almost every opened expedition is negative | | `bad_open_rate` | 0.905 | 0.908 | first-open quality does not improve | The open-quality metrics are generated by `src/coolrl_lost_cities/games/classic/deep_cfr/evaluate.py:124` and classify first-open actions around `src/coolrl_lost_cities/games/classic/deep_cfr/evaluate.py:546`. A "good" open has non-negative visible recoverable score; a "bad" open has no bonus path and negative recoverable score. Traversal gets deeper and slower late in training: - `traversal/nodes` mean rises from ~119k at iters 1-100 to ~199k at iters 1501-2000. - `traversal/avg_endpoint_depth` rises from ~212 to ~354. - `time/iteration_seconds` rises from ~12.0s to ~14.4s. - late node-limit cutoffs appear, but remain small on average: about 1.0% of endpoints in iters 1501-2000. This suggests the policy is changing game length and trajectory shape, but not toward higher-quality scoring. ## Interpretation The strongest signal is a current-vs-average policy split. Deep CFR trains advantage networks to define a current regret-matching policy, but the evaluated agent is the strategy network, which approximates the average policy. In this run, advantage loss improves while strategy loss and eval quality do not. That points to either: 1. the current policy is improving but the average strategy network or strategy memory is failing to represent it, or 2. the advantage networks are fitting targets that do not translate into better game play under the current traversal distribution. The current metrics cannot distinguish those two cases. The next diagnostic should evaluate the same checkpoint under both policy views: - current advantage networks + regret matching, - average strategy network. There is also a target-consistency issue worth auditing before a new long run. The recursive Cython traversal honors `outcome_unsampled_regret=zero` by setting unsampled legal targets to zero at `src/coolrl_lost_cities/games/classic/deep_cfr/traversal.pyx:930`. The interleaved traversal currently sets every legal target to `-node_value` before overwriting the sampled action at `src/coolrl_lost_cities/games/classic/deep_cfr/interleaved_traversal.py:386`. That makes the interleaved path behave like the `negative_node_value` variant even when the resolved config says `zero`. Theoretically, this is a state-wise constant shift and regret matching is invariant to action-independent constants. With function approximation and MSE, however, it changes training dynamics and should be treated as an experimental variable unless explicitly made consistent. ## Deep CFR Improvement Candidates ### 1. Evaluate current policy vs average policy first This is the highest-value diagnostic. If current regret-matching policy is materially stronger than the strategy network, focus on strategy memory, strategy weighting, and average-policy fitting. If both are weak, focus on advantage targets, exploration, and representation. ### 2. Make interleaved outcome targets match the config Either implement `outcome_unsampled_regret` in the interleaved scheduler or change the default/config documentation to state that interleaved outcome sampling uses the baseline-subtracted target. The former is cleaner because it lets `zero` vs `negative_node_value` become a controlled ablation. ### 3. Run a lower-variance traversal ablation Outcome sampling is correct but high variance for Lost Cities entry decisions. Before changing model size, test whether increasing `traversals_per_player` improves strict-heuristic score and expedition quality. A useful paired ablation would keep model/config fixed and change only traversal count, for example 280 vs 560 per player. ### 4. Add first-open regret diagnostics The current bad-open metric is outcome-facing and heuristic-visible. It tells us that opened expeditions are bad by recoverable score, but not whether the counterfactual regret target said the open action was bad at the decision point. Add diagnostics over advantage-memory targets by action class: - first-open action targets, - non-open alternatives in the same information state, - color/action class buckets, - target quantiles, not only means. If bad first-open actions have negative advantage targets but the policy still opens, suspect representation or strategy fitting. If their targets are not negative, suspect exploration distribution or the metric/objective mismatch. ### 5. Treat model-size scaling as a follow-up, not the first fix The earlier 1024x6 run showed a better 200-iteration signal than 512x3, but this 2000-iteration baseline shows a more basic policy-quality issue. Scaling capacity may help representation, but it will not tell us whether the average-policy training path is currently suppressing or distorting a better current policy. ## Recommended Next Plan 1. Implement or script current-vs-average evaluation for a saved checkpoint. 2. Run it on `iteration_00070.pt`, `iteration_00200.pt`, `iteration_01000.pt`, and `iteration_02000.pt` from the source run. 3. Audit and fix interleaved `outcome_unsampled_regret` handling. 4. Establish a new deterministic 512x3 baseline only after the target behavior is intentional and documented. 5. Then compare 512x3 vs 1024x6 under matched deterministic settings and W&B grouping. ## Practical Baseline Rule Do not promote `baseline-512x3-2000-dense-eval` as the canonical baseline. It is better classified as a diagnostic pre-determinism long run. The next canonical baseline should be rerun with: - deterministic traversal enabled, - `evaluation.eval_every=5`, - `checkpoint.save_every=50`, - explicit W&B group/job type, - and target semantics that match the resolved config.