Deep CFR gap 문서 최신화

This commit is contained in:
2026-05-06 23:37:43 +09:00
parent e802efabc3
commit 5a9166bf5a
+41 -83
View File
@@ -1,7 +1,8 @@
# Deep CFR v0 Gap vs Legacy coolrl # Deep CFR v0 Gap vs Legacy coolrl
This document compares the current `coolrl-lost-cities` Deep CFR v0 smoke This document compares the current `coolrl-lost-cities` Deep CFR v0
pipeline with the legacy Lost Cities Deep CFR implementation in `../coolrl`. implementation with the legacy Lost Cities Deep CFR implementation in
`../coolrl`.
The current implementation proves that Cython traversal primitives, PyTorch The current implementation proves that Cython traversal primitives, PyTorch
networks, memory collection, and a one-iteration smoke run can work together. It networks, memory collection, and a one-iteration smoke run can work together. It
@@ -15,59 +16,31 @@ Implemented:
- `random_rollout_value` - `random_rollout_value`
- `root_action_values` - `root_action_values`
- direct `GameState` C API use for legal actions and push/pop restoration - direct `GameState` C API use for legal actions and push/pop restoration
2. Minimal Cython information-state encoding. 2. Recursive Python Deep CFR traverser:
3. Small PyTorch MLP. - `traverse(state, traverser, iteration, depth)` logic
4. Simple in-memory sample storage. - terminal values
5. Minimal trainer that: - traverser vs opponent node behavior
- samples root action values with random rollouts - sampled action recursion
- builds advantage-like targets - sampled action value and node value calculation
- trains advantage networks and strategy network once - instantaneous regret collection at traverser nodes
6. Smoke tests for traversal restoration and trainer execution. - strategy-memory collection
- depth and node-budget cutoffs
3. Advantage-network-driven traversal policies:
- information-state encoding
- advantage network forward pass
- regret matching over legal actions
- sampled action recursion
4. Minimal Cython information-state encoding.
5. Small PyTorch MLP.
6. Simple in-memory sample storage with legal masks.
7. Legal-mask-aware advantage loss and masked strategy loss.
8. Smoke tests for traversal restoration and trainer execution.
This is a scaffold, not a complete Deep CFR algorithm. This is now a real single-process Deep CFR v0, but it is still not equivalent to
the legacy implementation.
## Major Algorithm Gaps ## Major Algorithm Gaps
### Recursive Deep CFR Traversal
Legacy `coolrl` has recursive outcome-sampling traversal. Current v0 does not.
Missing:
1. Recursive `traverse(state, traverser, iteration, depth)` logic.
2. Terminal value handling at every node.
3. Traverser node vs opponent node behavior.
4. Node value calculation from sampled actions.
5. Instantaneous regret calculation at traverser nodes.
6. Strategy-memory collection at traverser and/or opponent nodes.
7. Depth cutoff and node-budget cutoff.
This is the highest-priority gap.
### Network-driven Policies During Traversal
Current v0 does not use the advantage networks inside traversal. It estimates
root action values with random rollouts.
Legacy flow:
```text
encode information state
advantage network forward pass
regret matching over legal actions
sample action from policy
recurse
store regrets / strategy
```
Current flow:
```text
enumerate root actions
random rollout from each child
train on resulting root targets
```
### Outcome Sampling Controls ### Outcome Sampling Controls
Legacy `coolrl` supports: Legacy `coolrl` supports:
@@ -91,7 +64,8 @@ Legacy traversal supports:
4. rollout max-step timeouts 4. rollout max-step timeouts
5. cutoff stats 5. cutoff stats
Current v0 only uses random rollouts at the root-action helper level. Current v0 uses score-diff cutoff values for depth and node-budget cutoffs, but
does not support rollout-based cutoff values.
## Training and Memory Gaps ## Training and Memory Gaps
@@ -101,28 +75,11 @@ Legacy implementation has separate `AdvantageMemory` and `StrategyMemory` with:
1. capacity limits 1. capacity limits
2. reservoir sampling 2. reservoir sampling
3. legal masks stored with each sample 3. batch sampling
4. batch sampling 4. sample merging from traversal workers
5. sample merging from traversal workers
Current v0 stores simple `TrainingSample` objects in a list-like memory. Current v0 stores `TrainingSample` objects with legal masks, but still uses
list-like storage rather than true reservoir sampling.
### Legal-mask-aware Losses
Legacy advantage loss only trains legal action outputs:
```text
masked MSE over legal actions
```
Legacy strategy loss uses masked policy learning:
```text
masked logits -> log_softmax -> cross entropy against stored policy
```
Current v0 uses simple supervised MSE for both advantage and strategy targets.
This is enough for smoke testing, but not the intended training objective.
### Config System ### Config System
@@ -270,15 +227,16 @@ Current v0 has no league or checkpoint-snapshot opponent sampling.
Recommended implementation order: Recommended implementation order:
1. Implement real recursive Deep CFR traversal. 1. Add outcome-sampling controls.
2. Add legal-mask-aware advantage and strategy memories/losses. 2. Add rollout-based cutoff values.
3. Expand encoding to include public board, discard, and score features. 3. Replace list-like memory with reservoir memory.
4. Add checkpoint save/load. 4. Expand encoding to include public board, discard, and score features.
5. Add strategy-net bot adapter and evaluation integration. 5. Add checkpoint save/load.
6. Add CLI for train/eval/smoke. 6. Add strategy-net bot adapter and evaluation integration.
7. Add traversal stats and benchmark reporting. 7. Add CLI for train/eval/smoke.
8. Add multiprocessing workers only after the single-process algorithm is 8. Add traversal stats and benchmark reporting.
9. Add multiprocessing workers only after the single-process algorithm is
correct. correct.
The first three items are algorithm-critical. The rest are operationally useful The first four items are algorithm-critical. The rest are operationally useful
but should not block proving that the learning loop is correct. but should not block improving the learning loop.