Deep CFR legacy gap 문서 정리

This commit is contained in:
2026-05-07 00:02:25 +09:00
parent 46d841cf11
commit 573b59545f
+98 -219
View File
@@ -1,242 +1,121 @@
# Deep CFR v0 Gap vs Legacy coolrl
# Deep CFR v0 Status vs Legacy coolrl
This document compares the current `coolrl-lost-cities` Deep CFR v0
implementation with the legacy Lost Cities Deep CFR implementation in
`../coolrl`.
This document tracks the Lost Cities Deep CFR functionality in this repository
against the legacy implementation in `../coolrl`.
The current implementation proves that Cython traversal primitives, PyTorch
networks, memory collection, and a one-iteration smoke run can work together. It
is not yet equivalent to the legacy implementation.
## Implemented In This Repository
## Current v0
### Traversal
Implemented:
1. Cython traversal primitives:
- `random_rollout_value`
- `root_action_values`
- direct `GameState` C API use for legal actions and push/pop restoration
2. Recursive Python Deep CFR traverser:
- `traverse(state, traverser, iteration, depth)` logic
- terminal values
- traverser vs opponent node behavior
- sampled action recursion
- sampled action value and node value calculation
- instantaneous regret collection at traverser nodes
- strategy-memory collection
- depth and node-budget cutoffs
3. Advantage-network-driven traversal policies:
- information-state encoding
- advantage network forward pass
- regret matching over legal actions
- sampled action recursion
4. Minimal Cython information-state encoding.
5. Small PyTorch MLP.
6. Simple in-memory sample storage with legal masks.
7. Legal-mask-aware advantage loss and masked strategy loss.
8. Smoke tests for traversal restoration and trainer execution.
1. Recursive `traverse(state, traverser, iteration, depth)` logic.
2. Terminal value handling.
3. Traverser vs opponent node behavior.
4. Advantage-network-driven policies during traversal.
5. Regret matching over legal actions.
6. Sampled action recursion with node-value calculation.
7. Instantaneous regret collection at traverser nodes.
8. Strategy-memory collection.
9. Depth and node-budget cutoffs.
10. Outcome-sampling epsilon.
11. Sampled action probability correction.
12. Optional sampled value clipping.
13. Unsampled regret modes:
- `negative_node_value`
- `zero`
14. Score-diff and rollout-based cutoff values.
15. Random and safe-heuristic cutoff rollout policies.
16. Deck-draw chance sampling with state restoration.
This is now a real single-process Deep CFR v0, but it is still not equivalent to
the legacy implementation.
### Training And Memory
## Major Algorithm Gaps
Implemented:
### Outcome Sampling Controls
1. PyTorch advantage networks.
2. PyTorch strategy network.
3. Legal-mask-aware advantage loss.
4. Masked strategy cross-entropy loss.
5. Reservoir memory with capacity limits.
6. Batch sampling.
7. Legal masks stored with samples.
8. Single-process traversal.
9. Multiprocessing traversal worker batches.
10. Worker result merge in the parent trainer process.
Legacy `coolrl` supports:
### Encoding
1. `outcome_sampling_epsilon`
2. sampled action probability correction
3. optional sampled value clipping
4. unsampled regret modes:
- `negative_node_value`
- `zero`
Implemented information-state features:
Current v0 has none of these.
1. Phase flags.
2. Current player.
3. Encoded player.
4. Deck ratio.
5. Player hand slot features.
6. Public expedition summaries for both players.
7. Public discard summaries.
8. Public card counts.
9. Total score and score diff features.
10. Turn ratio.
11. Pending-discard one-hot.
12. Legal action mask.
### Cutoff Values
The encoding is still compact compared with the legacy feature set, but it now
contains the key public board, discard, score, and legal-action information.
Legacy traversal supports:
### Runtime Operations
1. score-diff cutoff values
2. random rollout cutoff values
3. safe-heuristic rollout cutoff values
4. rollout max-step timeouts
5. cutoff stats
Implemented:
Current v0 uses score-diff cutoff values for depth and node-budget cutoffs, but
does not support rollout-based cutoff values.
1. Checkpoint save/load.
2. Latest and per-iteration checkpoint files.
3. Config stored in checkpoints and `config.json`.
4. Strategy-net policy adapter.
5. Evaluation against registered classic bots.
6. Training CLI.
7. Evaluation CLI.
8. Traversal benchmark CLI.
9. Local run files:
- `config.json`
- `metrics.jsonl`
- `runtime_progress.json`
- `train.log`
10. Traversal benchmark metrics.
11. Self-play league snapshots.
12. Self-play league opponent selection from stored snapshots.
13. Safe-heuristic anchor opponent path.
## Training and Memory Gaps
## Remaining Differences From Legacy coolrl
### Reservoir Memory
The main Deep CFR v0 gaps listed earlier are now implemented. Remaining
differences are mostly experiment-system maturity and legacy-specific research
extras.
Legacy implementation has separate `AdvantageMemory` and `StrategyMemory` with:
Still smaller than legacy:
1. capacity limits
2. reservoir sampling
3. batch sampling
4. sample merging from traversal workers
1. Config is a single `DeepCFRConfig` dataclass rather than a deeply nested
YAML-first config tree.
2. Multiprocessing exists, but it is intentionally simple:
- no auto worker-count resolver
- no progress callback per worker batch
- no hotspot timing profile
3. Metrics logging exists, but no plotting/status command exists yet.
4. Checkpoint artifacts are local only; W&B artifact integration is not added.
5. Self-play league is simpler than legacy:
- snapshot sampling exists
- safe anchor path exists
- weighted recent/older/current buckets are not implemented
6. Legacy side experiments are not ported:
- imitation/pretraining commands
- policy-gradient fine-tuning commands
- legacy visualization helpers
Current v0 stores `TrainingSample` objects with legal masks, but still uses
list-like storage rather than true reservoir sampling.
These remaining items are not blockers for running and iterating on Deep CFR v0.
### Config System
## Suggested Next Steps
Legacy config is split into:
1. rules config
2. network config
3. encoding config
4. traversal config
5. optimization config
6. memory config
7. evaluation config
8. checkpoint config
9. run config
It also supports YAML loading and experiment-level overrides.
Current v0 has only `DeepCFRConfig`.
## Runtime and Operations Gaps
### Checkpointing
Legacy checkpoints include:
1. config
2. Lost Cities rules config
3. iteration
4. input/action dimensions
5. advantage networks
6. strategy network
7. optimizer states
8. self-play league snapshots
9. latest and per-iteration checkpoint files
Current v0 has no checkpoint save/load.
### Evaluation Integration
Legacy implementation can load a strategy checkpoint as a bot and evaluate it
against supported opponents.
Supported legacy eval flow includes:
1. random opponent
2. safe heuristic opponent
3. passive discard opponent
4. noisy/safe variants through the evaluation layer
5. many detailed gameplay metrics
Current v0 does not evaluate during training and does not expose a strategy-net
bot adapter.
### CLI
Legacy implementation has command-line tools for:
1. training
2. evaluation
3. evaluation suites
4. status/progress
5. plotting and visualization
6. traversal benchmarking
7. imitation/pretraining experiments
8. policy-gradient fine-tuning experiments
Current v0 has no CLI.
### Multiprocessing Workers
Legacy traversal can run through multiprocessing worker batches:
1. worker count resolution, including `auto`
2. traversal chunking
3. frozen network state dict transfer
4. worker-local traversal
5. result merging
6. progress logging
7. hotspot profiling
Current v0 runs in-process only.
### Metrics and Logging
Legacy training writes:
1. `metrics.jsonl`
2. `runtime_progress.json`
3. `train.log`
4. traversal nodes/sec
5. cutoff rates
6. endpoint depth buckets
7. advantage and strategy losses
8. evaluation metrics
9. hotspot timing metrics
Current v0 returns a small `IterationMetrics` object.
## Encoding Gaps
Legacy encoding is much richer. It includes:
1. phase one-hot
2. current-player indicator
3. player id
4. hand slots with card-type one-hot and empty-slot flag
5. both players' expedition card counts
6. expedition lengths
7. last numeric rank per expedition
8. discard pile card counts
9. discard pile length
10. discard pile top card
11. public card counts
12. deck ratio
13. turn-count ratio
14. pending-discard one-hot
15. optional derived playability features
16. optional slot-aware playability features
Current v0 includes only a minimal subset:
1. phase flags
2. current player
3. encoded player
4. deck ratio
5. player hand slot features
6. legal action mask
The current encoding should be expanded before serious training runs.
## Self-play League Gap
Legacy implementation supports self-play league opponent selection:
1. current networks
2. recent snapshots
3. older snapshots
4. safe-heuristic anchor
5. snapshot interval
6. maximum snapshot count
Current v0 has no league or checkpoint-snapshot opponent sampling.
## Practical Priority
Recommended implementation order:
1. Add outcome-sampling controls.
2. Add rollout-based cutoff values.
3. Replace list-like memory with reservoir memory.
4. Expand encoding to include public board, discard, and score features.
5. Add checkpoint save/load.
6. Add strategy-net bot adapter and evaluation integration.
7. Add CLI for train/eval/smoke.
8. Add traversal stats and benchmark reporting.
9. Add multiprocessing workers only after the single-process algorithm is
correct.
The first four items are algorithm-critical. The rest are operationally useful
but should not block improving the learning loop.
1. Add W&B/JSONL tracker abstraction on top of the existing local run files.
2. Add a status/plot command that reads `metrics.jsonl`.
3. Add richer benchmark output comparing single-process and multiprocessing.
4. Add weighted self-play league bucket selection if experiments need it.
5. Add W&B checkpoint artifacts after checkpoint quality is stable.