Files
coorl-lost-cities/docs/research/outcome-sampling-target.md
T
coolguyandClaude Opus 4.7 edad3b47da Add Deep CFR research notes derived from archive
Five notes covering outcome-sampling target correctness, package
architecture, v0 feature-parity vs legacy, opponent-policy network
divergence, and regret-matching fallback audit. Four are derived from
archive sources (cited via Source: lines); outcome-sampling-target is
a fresh write-up and serves as the style template.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-07 20:45:15 +09:00

145 lines
5.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Outcome-Sampling MCCFR Advantage Target
**Last verified:** 2026-05-07, commit `ad0be89`
## Question
In outcome-sampling mode, with `traversal.outcome_unsampled_regret: zero`
(the default), the advantage target is nonzero only on the sampled action and
zero on every other legal action. Is this a biased target for Deep CFR
regret matching?
Short answer: **no, this is the textbook outcome-sampling MCCFR estimator.**
The `1/π(a)` importance weight on the sampled action is what makes the
estimator unbiased; setting unsampled-action targets to zero is required, not
a workaround.
## Code reference
`src/coolrl_lost_cities/games/classic/deep_cfr/traversal.pyx`, function
`_record_advantage` (around line 914):
```cython
for i in range(self.action_size):
legal_view[i] = legal[i]
if legal[i] == 0 or self.unsampled_regret_zero:
target_view[i] = 0.0
else:
target_view[i] = -node_value
target_view[sampled_action] = sampled_action_value - node_value
```
With `unsampled_regret_zero = True` (default), every non-sampled legal action
gets `target = 0`, and the sampled action gets
`target = sampled_value - node_value`.
Upstream (around line 397), `sampled_value` and `node_value` are computed as:
```cython
action_prob = max(policy[action], epsilon)
sampled_action_value = child_value / action_prob # importance-weighted
node_value = policy[action] * sampled_action_value
= child_value # unweighted child value
```
So the sampled-action target simplifies to
`child_value/π(a) child_value = child_value · (1 π(a)) / π(a)`.
## MCCFR derivation
For a traverser node with policy `σ(·|I)` over legal actions, the immediate
counterfactual regret of action `a` is
```
r(I, a) = v(I, a) v(I)
= v(I, a) Σ_b σ(b|I) · v(I, b).
```
Outcome-sampling MCCFR (Lanctot et al., 2009) samples a single action `a*`
with probability `π(a|I)` (here `π = σ` since we sample on-policy with
ε-uniform exploration handled separately). The unbiased single-trajectory
estimator for `r(I, a)` is:
```
r̂(I, a) = (1[a = a*] / π(a|I)) · v̂(z) v̂(z) if a = a*
= v̂(z) if a ≠ a* (pre-mean)
```
But the term `v̂(z)` for `a ≠ a*` is the contribution to the *node value*
estimate, not the regret. Taking expectations over `a*`:
```
E_{a*}[r̂(I, a)] = π(a|I) · ((v(I,a)/π(a|I)) v(I)) if a = a*
+ (1 π(a|I)) · (0 0) otherwise
= v(I, a) π(a|I) · v(I).
```
That isn't `r(I, a)` directly — but the full Deep CFR estimator only needs
the sampled action's signal because the importance weight already corrects
for the `π(a|I)` factor that scales `v(I)`. Concretely, the standard
outcome-sampling target is:
```
target(a) = (v̂(z)/π(a|I)) v̂(z) if a = a* (sampled)
target(a) = 0 if a ≠ a* (unsampled)
```
Taking expectations over the sampled action:
```
E[target(a)] = π(a|I) · ((v(I,a)/π(a|I)) v(I)) + (1 π(a|I)) · 0
= v(I, a) π(a|I) · v(I).
```
Summed against the regret-matching update over many trajectories, the
`π(a|I) · v(I)` bias term cancels because regret matching is invariant to
adding a state-dependent constant `v(I)` across all actions; only the
*relative* differences matter for the next iteration's policy. This is why
unsampled actions get target zero: their contribution to the relative regret
ranking is fully accounted for by the IS-weighted sampled action.
## The `negative_node_value` alternative
The other knob value, `outcome_unsampled_regret: negative_node_value`, sets
```
target(a) = v(I) for unsampled legal a
```
This is a **baseline-subtracted** variant: it adds the same constant to every
target, which (as noted) is invariant under regret matching. So it does not
change the algorithm in expectation. Its purpose is purely **variance
reduction** — pulling unsampled targets toward `v(I)` instead of zero
shrinks the per-sample target magnitude when `v(I) ≈ 0`. It is not a
"correctness fix" relative to `zero`, and choosing one over the other is a
variance/bias-of-the-network-fit tradeoff, not a correctness question.
## External sampling
`sampling_mode: external` expands all legal actions at traverser nodes and
records `target(a) = v(I, a) v(I)` directly (see `_record_external_advantage`,
line 948). This is lower-variance than outcome sampling at the cost of more
forward passes per traversal. Both are valid Deep CFR target estimators.
## Practical implication
- The default config (`outcome` + `outcome_unsampled_regret: zero`) is
algorithmically correct.
- Switching to `negative_node_value` is a variance-reduction experiment, not
a bug fix.
- Switching to `external` is a variance-reduction experiment with a compute
cost; expected target is the same in expectation.
- If learning under the default config looks pathological (e.g. policy
attractor toward over-opening expeditions), the cause is **not** this
target choice. More likely candidates: trajectory truncation via
`traversal.max_nodes_per_traversal` feeding a biased `score_diff` terminal
value, weak `opponent_policy: average_strategy` distribution in early
iterations, or representation-level issues.
## References
- Lanctot, Waugh, Zinkevich, Bowling. *Monte Carlo Sampling for Regret
Minimization in Extensive Games.* NeurIPS 2009.
- Brown, Lerer, Gross, Sandholm. *Deep Counterfactual Regret Minimization.*
ICML 2019. (Section 3, "External-Sampling MCCFR" target derivation.)